MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
- 类型:arxiv
- 标识:2608.23035
- 链接:https://arxiv.org/abs/2608.23035
- 主分类:evaluation
- 形态:benchmark
- TLDR:As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot: GUI-centric benchmarks test surface-level screen manipulation while overlooking background tool use and long-horizon planning, whereas static function-calling benchmarks rely on offline API matching that is detached from real runtime constraints. To close this gap, we present MobilePA-Bench, an interactive, stateful, and tool-centric benchmark
- 副分类:agent
- 待LLM分类:否
- 标题中文:MobilePA-Bench:在复杂真实任务上对移动端 Planner Agent 的基准评测
- TLDR中文:随着端侧 LLM Agent 演变为个人副驾驶,移动操作系统已成为该范式的关键试验场,亟需严格的能力评测。然而现有基准可分为两类,各自存在关键盲区:以 GUI 为中心的基准仅测试表层屏幕操作,忽略了后台工具调用与长程规划;而静态 function-calling 基准依赖离线 API 匹配,与真实运行时约束脱节。为弥合这一差距,我们提出 MobilePA-Bench,一个交互式、有状态、以工具为中心的基准。
- 来源文件:
- /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json