RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

  • 类型:arxiv
  • 标识:2609.25636
  • 链接:https://arxiv.org/abs/2609.25636
  • 主分类:agent
  • 形态:benchmark
  • TLDR:Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score highly while barely using it. We introduce RoboFollow, a diagnostic benchmark with three principles: (1) High Scene Entropy: each training scene supports multiple kinematically distinct task branches, making vision alone insufficient and forcing reliance on language. (2) Hierarchi
  • 待LLM分类:否
  • 标题中文:RoboFollow:揭示具身智能体中指令遵循的幻象
  • TLDR中文:现代具身智能体取得了很高的成功率,但其真实指令遵循能力远弱于这些数字所暗示的水平。我们将这种假象追溯到一种结构性现象——低场景熵:当视觉场景仅容许一个有效任务时,语言变得冗余,策略即便几乎不使用语言也能取得高分。我们提出 RoboFollow,一个诊断式 benchmark,遵循三条原则:(1) 高场景熵:每个训练场景支持多个运动学上不同的任务分支,使仅靠视觉不足,迫使模型依赖语言。(2) 层次
  • 来源文件
  • /inbox/tom/_candidates/2026-09-23-agent-rag-longcontext-candidates.json