MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

  • 类型:arxiv
  • 标识:2608.23035
  • 链接:https://arxiv.org/abs/2608.23035
  • 主分类:evaluation
  • 形态:benchmark
  • TLDR:As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot: GUI-centric benchmarks test surface-level screen manipulation while overlooking background tool use and long-horizon planning, whereas static function-calling benchmarks rely on offline API matching that is detached from real runtime constraints. To close this gap, we present MobilePA-Bench, an interactive, stateful, and tool-centric benchmark
  • 副分类:agent
  • 待LLM分类:否
  • 标题中文:MobilePA-Bench:在复杂真实任务上对移动端 Planner Agent 的基准评测
  • TLDR中文:随着端侧 LLM Agent 演变为个人副驾驶,移动操作系统已成为该范式的关键试验场,亟需严格的能力评测。然而现有基准可分为两类,各自存在关键盲区:以 GUI 为中心的基准仅测试表层屏幕操作,忽略了后台工具调用与长程规划;而静态 function-calling 基准依赖离线 API 匹配,与真实运行时约束脱节。为弥合这一差距,我们提出 MobilePA-Bench,一个交互式、有状态、以工具为中心的基准。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json