Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
- 类型:arxiv
- 标识:2609.35767
- 链接:https://arxiv.org/abs/2609.35767
- 主分类:multimodal
- 形态:method
- TLDR:Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Refl
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-29-agent-rag-longcontext-candidates.json