PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

  • 类型:arxiv
  • 标识:2609.38597
  • 链接:https://arxiv.org/abs/2609.38597
  • 主分类:multimodal
  • 形态:method
  • TLDR:Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding
  • 待LLM分类:否
  • 标题中文:PixelUMM:无编码器的统一图像与视频理解与生成
  • TLDR中文:统一多模态模型(UMM)通常依赖独立的视觉表示分别完成理解与生成,这增加了视觉上下文长度,并使其难以与既有视觉-语言预训练流程集成。近期 pixel-space modeling 的进展提供了一种无编码器的替代方案,但将该范式从图像扩展到视频并非易事:视频理解与生成采用不同的时间表示,统一视觉接口的设计仍是开放问题。本文提出 PixelUMM,一种用于统一图像与视频理解的无编码器模型……
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-02-agent-rag-longcontext-candidates.json