OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
- 类型:arxiv
- 标识:2609.21465
- 链接:https://arxiv.org/abs/2609.21465
- 主分类:multimodal
- 形态:benchmark
- TLDR:We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two constraints: data availability and evaluation. Recordings of people using their own devices are scarce.
- 副分类:evaluation
- 待LLM分类:否
- 标题中文:OmniVChat:面向原生音视频对话的合成、基准与训练
- TLDR中文:我们将 OmniVChat(Omni Video Chat)定义为用户与 omni 模型之间的原生音视频对话任务。在 OmniVChat 中,omni 模型直接并同时接收用户的音频与视频,并返回文本。用户查询嵌入在音频与视频中,无需独立的文本问题、外部 caption 或语音识别。直接的音视频输入降低了外部延迟与计算开销,同时保留感知线索。然而,OmniVChat 研究受限于两方面:数据可用性与评测。用户使用自有设备录制的素材稀缺。
- 来源文件:
- /inbox/tom/_candidates/2026-09-21-agent-rag-longcontext-candidates.json