MOSS-VL Technical Report

  • 类型:arxiv
  • 标识:2608.15045
  • 链接:https://arxiv.org/abs/2608.15045
  • 主分类:multimodal
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar + OpenAlex
  • S2被引:0
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay silent, and when to revise; and a staged curriculum concentrates all real-time-specific training in one light final stage over a strong offline foundation. Offline, MOSS-VL-Instruct is competitive at
  • OpenAlex ID:W7203670747
  • OpenAlex DOI:10.48550/arxiv.2608.15045
  • DOI:10.48550/arxiv.2608.15045
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://doi.org/10.48550/arxiv.2608.15045
  • OpenAlex更新:2026-08-23
  • 待LLM分类:否
  • 标题中文:MOSS-VL 技术报告
  • TLDR中文:我们提出 MOSS-VL,一个开源视觉-语言模型系列,将实时交互(边说边看)视为一等能力。它贯穿整个栈进行协同设计:语言解码器仅通过门控交叉注意力访问视觉,因此模型在生成过程中可以自然地感知新输入帧;合成的交互语料用于监督何时说话、何时沉默、何时修正;分阶段课程将所有实时相关训练集中在一个轻量的最终阶段,基于强大的离线基础模型。在离线场景下,MOSS-VL-Instruct 在
  • 来源文件
  • /inbox/tom/_candidates/2026-08-18-agent-rag-longcontext-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]