RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents
- 类型:arxiv
- 标识:2609.25636
- 链接:https://arxiv.org/abs/2609.25636
- 主分类:agent
- 形态:benchmark
- TLDR:Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score highly while barely using it. We introduce RoboFollow, a diagnostic benchmark with three principles: (1) High Scene Entropy: each training scene supports multiple kinematically distinct task branches, making vision alone insufficient and forcing reliance on language. (2) Hierarchi
- 待LLM分类:否
- 标题中文:RoboFollow:揭示具身智能体中指令遵循的幻象
- TLDR中文:现代具身智能体取得了很高的成功率,但其真实指令遵循能力远弱于这些数字所暗示的水平。我们将这种假象追溯到一种结构性现象——低场景熵:当视觉场景仅容许一个有效任务时,语言变得冗余,策略即便几乎不使用语言也能取得高分。我们提出 RoboFollow,一个诊断式 benchmark,遵循三条原则:(1) 高场景熵:每个训练场景支持多个运动学上不同的任务分支,使仅靠视觉不足,迫使模型依赖语言。(2) 层次
- 来源文件:
- /inbox/tom/_candidates/2026-09-23-agent-rag-longcontext-candidates.json