MOSS-VL Technical Report
- 类型:arxiv
- 标识:2608.15045
- 链接:https://arxiv.org/abs/2608.15045
- 主分类:multimodal
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar + OpenAlex
- S2被引:0
- OpenAlex被引:0
- 影响力被引:0
- TLDR:We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay silent, and when to revise; and a staged curriculum concentrates all real-time-specific training in one light final stage over a strong offline foundation. Offline, MOSS-VL-Instruct is competitive at
- OpenAlex ID:W7203670747
- OpenAlex DOI:10.48550/arxiv.2608.15045
- DOI:10.48550/arxiv.2608.15045
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://doi.org/10.48550/arxiv.2608.15045
- OpenAlex更新:2026-08-23
- 待LLM分类:否
- 标题中文:MOSS-VL 技术报告
- TLDR中文:我们提出 MOSS-VL,一个开源视觉-语言模型系列,将实时交互(边说边看)视为一等能力。它贯穿整个栈进行协同设计:语言解码器仅通过门控交叉注意力访问视觉,因此模型在生成过程中可以自然地感知新输入帧;合成的交互语料用于监督何时说话、何时沉默、何时修正;分阶段课程将所有实时相关训练集中在一个轻量的最终阶段,基于强大的离线基础模型。在离线场景下,MOSS-VL-Instruct 在
- 来源文件:
- /inbox/tom/_candidates/2026-08-18-agent-rag-longcontext-candidates.json
- [S2 enrich]
- [OpenAlex backfill]