Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

  • 类型:arxiv
  • 标识:2609.09143
  • 链接:https://arxiv.org/abs/2609.09143
  • 主分类:multimodal
  • 形态:method
  • TLDR:Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction. We examine how these losses scale and relate to downstream performance, then use them to study multimodal learnability--
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:在统一多模态模型中将图像 tokenizer 视为视觉语言的研究
  • TLDR中文:图像 tokenizer 定义了统一多模态模型的"视觉语言",但其研究通常依赖孤立指标或仅针对生成/理解的评估。这些评估无法充分刻画视觉 token 与文本联合建模时的行为。我们构建了受控的纯自回归测试平台,在多模态持续预训练过程中追踪文本、图像、文生图(T2I)与图生文(I2T)预测的任务专属验证损失。我们考察这些损失的缩放特性及其与下游性能的关系,并据此研究多模态可学习性……
  • 来源文件
  • /inbox/tom/_candidates/2026-09-12-agent-rag-longcontext-candidates.json