Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

  • 类型:arxiv
  • 标识:2609.35767
  • 链接:https://arxiv.org/abs/2609.35767
  • 主分类:multimodal
  • 形态:method
  • TLDR:Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Refl
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-09-29-agent-rag-longcontext-candidates.json