Reasoning with Image Generation
- 类型:arxiv
- 标识:2609.16409
- 链接:https://arxiv.org/abs/2609.16409
- 主分类:multimodal
- 形态:method
- TLDR:Chain-of-thought reasoning has revolutionized natural language processing by enabling large language models (LLMs) to decompose problems into intermediate steps before answering. Yet confining reasoning to the textual domain presents limitations for tasks requiring direct manipulation of visual representations. Recent efforts augment multimodal LLMs with external visual expert tools such as depth estimation or object detection modules, but these remain fundamentally limited by their reliance on narrow, rigid operations that cannot flexibly generate or transform visual content. We propose ReIma
- 待LLM分类:否
- 标题中文:结合图像生成的推理
- TLDR中文:思维链推理通过让 LLM 在回答前将问题分解为中间步骤,彻底改变了自然语言处理。然而,将推理局限于文本域,对需要直接操作视觉表示的任务存在局限。近期工作通过为多模态 LLM 引入深度估计、目标检测等外部视觉专家工具来增强能力,但仍受限于这些操作固有的狭窄与刚性,无法灵活生成或变换视觉内容。我们提出 ReIma
- 来源文件:
- /inbox/tom/_candidates/2026-09-30-agent-rag-longcontext-candidates.json