UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
- 类型:arxiv
- 标识:2608.08676
- 链接:https://arxiv.org/abs/2608.08676
- 主分类:multimodal
- 形态:method
- TLDR:Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine-grained visual details, leading to poor pixel reconstruction and limiting their use in reconstruction-sensitive tasks such as image generation and editing. In this work, we ask whether understanding, generation, and editing can be modeled in a single visual representation space built from a pretrained semantic ViT. We show that the frozen Transformer blocks of a semantic ViT are not intrinsically unable to preserve
- 待LLM分类:否
- 标题中文:UniSpace:统一视觉表征与可扩展多模态建模
- TLDR中文:语义视觉编码器已成为多模态理解与图像生成中语义条件的关键视觉接口。然而其最终 token 丢弃了细粒度视觉细节,导致像素重建质量较差,限制了其在图像生成与编辑等对重建敏感的任务中的应用。本工作探讨理解、生成与编辑能否在由预训练语义 ViT 构建的单一视觉表征空间中建模。我们证明,语义 ViT 的冻结 Transformer 块本身并非无法保留
- 来源文件:
- /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json