Hugging Face Daily Papers · · 3 min read

VibeVoice-ASR-Streaming Technical Report

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

VibeVoice-ASR-Streaming, a unified streaming ASR model that continuously transcribes ''who said what'' as speech arrives,</p>\n","updatedAt":"2026-09-03T05:32:54.312Z","author":{"_id":"646c408336505117e22f4b36","avatarUrl":"/avatars/8201da3c4ec51c8e113446f9578d44f6.svg","fullname":"zhiliang","name":"zzliang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9596290588378906},"editors":["zzliang"],"editorAvatarUrls":["/avatars/8201da3c4ec51c8e113446f9578d44f6.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.02812","authors":[{"_id":"6a990641fea818274321ff9f","name":"Yujie Tu","hidden":false},{"_id":"6a990641fea818274321ffa0","name":"Zhiliang Peng","hidden":false},{"_id":"6a990641fea818274321ffa1","name":"Jianwei Yu","hidden":false},{"_id":"6a990641fea818274321ffa2","name":"Li Dong","hidden":false},{"_id":"6a990641fea818274321ffa3","name":"Songchen Xu","hidden":false},{"_id":"6a990641fea818274321ffa4","name":"Yaoyao Chang","hidden":false},{"_id":"6a990641fea818274321ffa5","name":"Wenhui Wang","hidden":false},{"_id":"6a990641fea818274321ffa6","name":"Zilong Wang","hidden":false},{"_id":"6a990641fea818274321ffa7","name":"Zehua Wang","hidden":false},{"_id":"6a990641fea818274321ffa8","name":"Yan Xia","hidden":false},{"_id":"6a990641fea818274321ffa9","name":"Jiajun Zhang","hidden":false},{"_id":"6a990641fea818274321ffaa","name":"Xie Chen","hidden":false},{"_id":"6a990641fea818274321ffab","name":"Furu Wei","hidden":false}],"publishedAt":"2026-09-02T00:00:00.000Z","submittedOnDailyAt":"2026-09-03T00:00:00.000Z","title":"VibeVoice-ASR-Streaming Technical Report","submittedOnDailyBy":{"_id":"646c408336505117e22f4b36","avatarUrl":"/avatars/8201da3c4ec51c8e113446f9578d44f6.svg","isPro":false,"fullname":"zhiliang","user":"zzliang","type":"user","name":"zzliang"},"summary":"Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ''who said what'' as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.","upvotes":8,"discussionId":"6a990641fea818274321ffac","ai_summary":"A streaming, LLM-based end-to-end model unifies speaker-attributed speech recognition and diarization for low-latency real-time applications.","ai_keywords":["speaker-attributed ASR","end-to-end","LLM-based","streaming","speaker diarization","lookahead audio","WER/CER"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"5e6485f787403103f9f1055e","name":"microsoft","fullname":"Microsoft","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583646260758-5e64858c87403103f9f1055d.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"646c408336505117e22f4b36","avatarUrl":"/avatars/8201da3c4ec51c8e113446f9578d44f6.svg","isPro":false,"fullname":"zhiliang","user":"zzliang","type":"user"},{"_id":"68ac9da42166e79cec1b1571","avatarUrl":"/avatars/e2e138a8df6026e6f26df54ec32ed47a.svg","isPro":false,"fullname":"hulu","user":"hululuhu","type":"user"},{"_id":"6a97f937032af4b22d34bb54","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/xsK8NqdfHCu_IGRjdsb4z.jpeg","isPro":false,"fullname":"Zhiliang Peng","user":"zhiliangp","type":"user"},{"_id":"67ecd6178647cfa1775f75ed","avatarUrl":"/avatars/98882cc58dc0a5de94df765d523d92c9.svg","isPro":false,"fullname":"Furu Wei","user":"frontierai","type":"user"},{"_id":"69ccb304d5dcac9a4ac1772d","avatarUrl":"/avatars/826ea095e4b0eef72e147e10d1d06c39.svg","isPro":false,"fullname":"Isabella Wright","user":"victoriapqj64","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"67320e4996d5da4801a69199","avatarUrl":"/avatars/bf85472384886c19121a5b32bb4dbeea.svg","isPro":false,"fullname":"AlexTYJ","user":"AlexTYJ","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"5e6485f787403103f9f1055e","name":"microsoft","fullname":"Microsoft","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583646260758-5e64858c87403103f9f1055d.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.02812.md","query":{}}">
Papers
arxiv:2609.02812

VibeVoice-ASR-Streaming Technical Report

Published on Sep 2
· Submitted by
zhiliang
on Sep 3
Authors:
,

Abstract

A streaming, LLM-based end-to-end model unifies speaker-attributed speech recognition and diarization for low-latency real-time applications.

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ''who said what'' as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.

Community

Paper submitter about 3 hours ago

VibeVoice-ASR-Streaming, a unified streaming ASR model that continuously transcribes ''who said what'' as speech arrives,

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.02812
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.02812 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers