Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning

  • 类型:arxiv
  • 标识:2608.26993
  • 链接:https://arxiv.org/abs/2608.26993
  • 主分类:multimodal
  • 形态:method
  • TLDR:Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an image editor can faithfully realize the required transformation. We introduce Aphanta, an automated task-discovery and closed-loop diagnostic framework for the MLLM -> image editor -> MLLM pipeline. Aphanta evaluates three conditions---direct reasoning, reasoning with an editor-generated intermediate, and reasoning with an idealized reference intermediate---to separate potential visual headroom from the practical utility
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-08-28-agent-rag-longcontext-candidates.json