VibeVoice-ASR-Streaming Technical Report
- 类型:arxiv
- 标识:2609.02812
- 链接:https://arxiv.org/abs/2609.02812
- 主分类:multimodal
- 形态:method
- TLDR:Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and pr
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-03-agent-rag-longcontext-candidates.json