Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

  • 类型:arxiv
  • 标识:2607.24904
  • 链接:https://arxiv.org/abs/2607.24904
  • 主分类:multimodal
  • 形态:method
  • 被引:1
  • 被引来源:Semantic Scholar
  • S2被引:1
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:Mage-VL is presented, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction and establishes AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes.
  • OpenAlex ID:W7171699370
  • OpenAlex DOI:10.48550/arxiv.2607.24904
  • DOI:10.48550/arxiv.2607.24904
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://doi.org/10.48550/arxiv.2607.24904
  • OpenAlex更新:2026-08-25
  • 待LLM分类:否
  • 标题中文:Mage-VL:面向高效编解码器原生流式多模态的基础模型
  • TLDR中文:本文提出 Mage-VL,一种面向实时多模态理解与交互的高效 codec-native 流式基础模型,并构建了 AI4AI 数据流水线,涵盖面向多模态 captioning 的 prompt-code 联合优化与以 AI 驱动的性能诊断,以指导训练方案。
  • 来源文件
  • /inbox/tom/_candidates/2026-07-29-agent-rag-longcontext-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]