VibeVoice-ASR Ships 1.5B and 7B Open Weights for Streaming Speaker-Attributed Speech Recognition
Summary
A new arXiv technical report unveils VibeVoice-ASR-Streaming, an LLM-based system that unifies transcription and speaker identification in one pass by interleaving fixed-size audio chunks with lookahead audio and prior text. The authors released 1.5B and 7B model weights plus inference code; the 7B model records the lowest average WER/CER across five evaluation sets and best or tied-best speaker attribution on 12 of 13 settings, eliminating the need for a separate diarization stack.
Originally reported by huggingface.co
Read the original article →Original headline: VibeVoice-ASR Ships 1.5B and 7B Open Weights for Streaming Speaker-Attributed Speech Recognition