arXiv:2609.38597 · 多模态
PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
PixelUMM:无编码器的统一图像与视频理解与生成
PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
- 类型:arxiv
- 标识:2609.38597
- 链接:https://arxiv.org/abs/2609.38597
- 主分类:multimodal
- 形态:method
- TLDR:Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding
- 待LLM分类:否
- 标题中文:PixelUMM:无编码器的统一图像与视频理解与生成
- TLDR中文:统一多模态模型(UMM)通常依赖独立的视觉表示分别完成理解与生成,这增加了视觉上下文长度,并使其难以与既有视觉-语言预训练流程集成。近期 pixel-space modeling 的进展提供了一种无编码器的替代方案,但将该范式从图像扩展到视频并非易事:视频理解与生成采用不同的时间表示,统一视觉接口的设计仍是开放问题。本文提出 PixelUMM,一种用于统一图像与视频理解的无编码器模型……
- 来源文件:
- /inbox/tom/_candidates/2026-10-02-agent-rag-longcontext-candidates.json