本文提出 π-Bench,一个用于评估主动式协助能力的基准,包含跨 5 个领域特定用户画像的 100 个多轮任务,用于评估 Agent 在长交互中预见并满足用户需求的能力,联合衡量长周期轨迹中的主动性与任务完成度,更贴近真实使用场景。$-Bench is introduced, a benchmark for proactive assistance comprising 100 multi-turn tasks across 5 domain-specific user personas that evaluates agents'ability to anticipate and address user needs over extended interactions, jointly measuring proactivity and task completion in long-horizon trajectories that better reflect real-world use.
论文
66 张论文卡片 · Agent 智能体 · 评测集
Agents' Last Exam(ALE)是一个面向长时序、具有经济价值且结果可验证的真实任务的 AI Agent 评测基准,旨在弥合基准测试表现与 GDP 相关影响之间的差距,而非仅仅作为排行榜。Agents'Last Exam (ALE) is introduced, a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes, intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP relevant impact.
本文构建一种针对 GitHub issue 的持久会话评估方法,锚定在单一 base commit,对线性顺序探索与非线性、领域范围的并行 agentic 探索进行比较。This work constructs an approach for persistent-session evaluation of GitHub issues anchored at a single base commit, and compares linear sequential exploration against non-linear, domain-scoped parallel agentic exploration.
结果表明,在所评估的协调者–工作者设定下,多 Agent 系统中的隐私风险主要由架构层面的协调通道决定,而非仅取决于最终输出行为:风险来源于对标准输出级防御不可见的内部通道。Results suggest, within the evaluated coordinator-worker setting, that privacy risk in multi-agent systems is strongly shaped by architectural coordination channels rather than final-output behavior alone: it arises from internal channels that remain invisible to standard output-level defenses.
介绍 MCP-Persona,这是首个专为评估 Agent 在真实场景、个性化 MCP 工具上的表现而设计的基准,并揭示了当前 Agent 在个性化工具使用上的显著不足,从而凸显该基准在发现并解决这些局限上的关键作用。MCP-Persona is introduced, the first benchmark specifically designed for evaluating agent performance on real-world, personalized MCP tools and demonstrates their significant struggles with personalized tool use, thereby highlighting the benchmark's crucial role in identifying and addressing these limitations.
本文设计了一项对比研究,结合受控定量实验与配对轨迹分析,并将观察结果归纳为一个包含三个高层类别与十二种 Skill 使用模式的分类法,表明当噪声轨迹转化为稳定执行的过程性锚点时,Skill 便会发挥作用。This work designs a contrastive study that combines controlled quantitative experiments with paired trajectory analysis and consolidates observations into a taxonomy of three high-level categories and twelve skill-use modes, showing that skills work when noisy trajectories become procedural anchors that stabilize execution.
一项配对消融实验在保留仓库与可执行工程上下文的同时移除显式科学指导,表明科学知识并非一律有益:可靠信息能约束修复、提升平均表现与 token 效率,而错位指导则会诱发锚定,不必然提升精确修复成功率。A paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context shows that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success.
在所测试的同伴排序信息流下,鲁棒的结论是词法层面的趋同,而非对一般意见的捕获,也不是合成 LLM-Agent 群体中的一般性协调优势;该结果未估计其对人类或真实生产平台的影响。The robust result is lexical convergence under the tested peer-ranked feed, not general opinion capture or a general coordination advantage in the synthetic LLM-agent populations; it does not estimate effects on people or production platforms.
本文提出 GameXpert-Bench,将游戏开发的三个生命周期阶段以 coding agent 操作化为三条互补的基准轨道,并发现现有 agent 在生成可玩基础框架和实现显式需求方面更可靠,而在发现缺陷、验证运行时行为以及跨变更保持功能一致性方面能力较弱。GameXpert-Bench is introduced, which operationalizes the three lifecycle stages of game development with a coding agent as three complementary benchmark tracks, and finds current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.
本文发布 SWE Refactor Bench,一个包含 20 项全仓库迁移的基准,涵盖 4 类技术债务,作为开发面向可靠全仓库迁移的编码 Agent 的严格测试平台。SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.
提出 EASEL benchmark,用于评估受控的灵巧视觉工具使用,以参考引导的视觉重建为主代理任务:agent 逐步在画布上绘制以匹配参考图像。EASEL is proposed, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image.
提出 LoopArena,一个用于评估一个模型在长时任务中引导另一个独立 coding agent 能力的 benchmark,并在执行范围与成本各不相同的三个互补设置下评估该能力。This work introduces LoopArena, a benchmark for evaluating how well one model can guide a separate coding agent through a long-running task, and evaluates this ability in three complementary settings that differ in execution scope and cost.
提出 MnIST-PRO 基准,将 Agent 感知能力隔离测试:把 MNIST 数字识别转换为带回看约束的序列式 glimpse 搜索任务,并评估了十个多模态模型,结果表明仅获取视觉证据是不够的,Agent 还必须能够构建并更新可靠的感知状态。MnIST-PRO is addressed, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints, and ten multimodal models are evaluated, showing that simply acquiring visual evidence is not enough and agents must also be able to build and update a reliable perceptual state.
本文提出 UI-Venus-2,一个面向移动、Web 和桌面环境的通用 GUI 基础 Agent,采用统一的闭环推理-行动框架,并通过集成安全感知机制来确保关键操作的可控执行。UI-Venus-2 is presented, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework that integrates safety-aware mechanisms to ensure controlled execution of consequential actions.
本文提出 agentic data cracking,将非结构化数据自适应、推测性地结构化为推理自身的副产品,是面向非结构化数据上 Agentic reasoning 的下一代数据基础设施的第一步。This work proposes agentic data cracking, a method that structures unstructured data adaptively and speculatively as a byproduct of reasoning itself, a first step toward next-generation data infrastructure for agentic reasoning over unstructured data.
本文提出一种知识门控的任务构建协议,将任务指令与一个包含私有约定、参考表与效用算子的紧凑工件分离,并证明被保留的任务能够改善 post-training。A knowledge-gated task-construction protocol is introduced that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators, and it is shown that the retained tasks improve post-training.
提出了 τ^τ-bench(读作 hyper-tau-bench),一个将 Agent 构建作为任务的 benchmark,将协作式 Agent 构建工作转化为面向 coding agent 的可度量目标。The $\tau^\tau$-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task, is introduced to turn the work of cooperative agent building into a measurable target for coding agents.
HarvestBench 是首个为避免 side effect 标定价格并将 side effect 命名为生物的 benchmark;在六个模型中有四个的每次作答遭遇 kill rate 对价格变化敏感。HarvestBench is the first benchmark to put a price on avoiding a side effect and name the side effect as a living creature and four out of six models'kill rate per answered encounter were sensitive to price changes.
提出一个自演化 Agent 框架,通过识别 Agent 轨迹中不稳定、低一致性的步骤并将其转化为情节记忆,以供后续运行调用,从而缩小一致性 gap。This work presents a self-evolving agent framework that reduces the consistency gap by identifying unstable, low-consistency steps in agent trajectories and converting them into episodic memory the agent can draw on in future runs.
推出 Atria Dawn Preview,一个面向科学研究与工程工作流的基础 Agent 语言模型,旨在拓展真实场景下 Agent 生产力的前沿。Atria Dawn Preview is introduced, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world.
该工作提出 PACT(Pressure-Applied Compliance Testing),一个用于评估 AI agent 在压力下遵守规则的基准,覆盖十二个受监管的企业领域与四十八个场景,每个场景均为真实的多轮对话。This work introduces PACT (Pressure-Applied Compliance Testing), a benchmark for rule-following under pressure in AI agents assisting employees in daily tasks across twelve regulated enterprise domains and forty-eight scenarios, each set in a realistic multi-turn conversation.
APort Vault 是面向工具使用 AI Agent 支付授权的基准。它在公开 CTF 活动中重放了人类编写的 4,371 次针对真实支付 Agent 的攻击,覆盖 8 家实验室的 14 个模型、五种策略配置与两条重放轨道,每条轨道分别在有/无确定性的 pre-action check(实现 Open Agent Passport (OAP) 规范)下执行,共计完成 225,964 次评测。我们每次评测报告五个独立事件,因为将它们合并正是 Agent 基准产生无法经得起审查的数字的方式。请求很常见,且其速率差异巨大APort Vault is a benchmark for payment authorization in tool-using AI agents. It replays 4,371 attacks written by humans against a live payment agent during a public capture-the-flag event, across 14 models from 8 labs, five policy configurations and two replay tracks, with and without a deterministic pre-action check implementing the Open Agent Passport (OAP) specification. 225,964 evaluations completed. We report five distinct events per evaluation, because collapsing them is how an agent benchmark produces a number that does not survive review. Requests are common and their rate differs far
构建 Taste-Bench,一个由 Agent 在工程与研究任务中产生的轨迹自动构建的品味问题基准,并证明品味是可训练的。Taste-Bench is built, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks, and it is shown that taste can be trained.
本文提出 AgenticRAGTracer,这是首个主要由大语言模型自动构建、专为支持逐步验证而设计的 Agentic RAG 基准。AgenticRAGTracer is introduced, the first Agentic RAG benchmark that is primarily constructed automatically by large language models and designed to support step-by-step validation, and is primarily constructed automatically by large language models and designed to support step-by-step validation.
本文提出了 RoboFollow,一个基于三条原则的诊断 benchmark,揭示了真正的指令跟随能力是具身 Agent 在执行指令时被忽视的关键瓶颈。This work introduces RoboFollow, a diagnostic benchmark with three principles, which exposes genuine instruction following as a critical, overlooked bottleneck in embodied agents' ability to follow instructions.
EMBODIEDSWE-GEN 将 coding agent 的单一解决方案扩展为大规模多样化轨迹用于训练 VLA,并表明仅在 coding agent 生成的仿真演示上微调的 VLA,即可在真实机器人上完成长时任务。EMBODIEDSWE-GEN, which expands a single solution from coding agent into large diverse trajectories for training a VLA, and shows that a VLA finetuned solely on coding-agent-generated simulation demonstrations completes a long-horizon task on real robot.
为在传统二元任务成功之外量化协作有效性,本文提出 Causal Collaboration Effectiveness (CCE),一种基于图的指标,用于追踪 agent 动作之间的因果依赖,并度量团队投入中实际促成最终结果的比例。To quantify collaboration effectiveness in addition to conventional binary task success, Causal Collaboration Effectiveness (CCE) is proposed, a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome.
本文提出 LibraryDesignBench,这是一个两阶段基准,Agent 根据一份定义所需能力和潜在用例但不规定具体设计的规范,实现一个功能完备的 library。LibraryDesignBench is introduced, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design.
本评估刻画了前沿 agent 如何结合源代码级执行、应用截图与图形交互来生成经过验证的软件变更,考察了跨领域与不同任务信息需求下的表现,以及与成功修复相关的开发行为。This evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes, and examines performance across domains and task information requirements, alongside the development behaviors associated with successful repairs.
本文提出 BIABench,一个由 16 个从已发表生物研究重建的任务组成的基准,保留了其科学问题、成像数据与真值标注,为评估并最终训练可靠的、面向长程的生物图像分析 agent 提供可验证的框架。BIABench is introduced, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth, and provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.
提出了 DyadMem,一个双领域、全流程的记忆基准,伴随大量标注工作,并提出了新定义——用户条件关系型智能体记忆(URAM),以推动该领域发展。DyadMem is a dual-domain, full-pipeline memory benchmark with extensive annotation efforts for advancing the domain's development and is proposed with the proposed new definition User-conditioned Relational Agent Memory (URAM).
提出 UndoBench,一个覆盖 8 个企业领域、36 个基础工作流与 36 个故障场景的基准测试,通过相同种子下的反事实配对试验以及链路级效应历史与环境状态预言机,将任务能力与恢复能力解耦。UndoBench is introduced, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles.
黑盒优化(BBO)出现在许多目标和工程问题中,其目标函数评估代价高昂且次数有限。近期的大语言模型(LLM)智能体通过结合任务语义、计算、优化工具以及反馈驱动的决策,提供了一种新的 BBO 解决方式,因与数学严谨工具的集成而展现出巨大潜力。然而,现有的 Agentic BBO 研究使用不同的任务领域和系统配置,导致结果难以比较,且单个设计选择的影响难以孤立分析。因此,我们引入Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools. However, existing agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We therefore i
本工作将带外防御组织为经典完整性保护、引用监控与最小权限的具体实例,对它们覆盖与未覆盖的内容进行结构化对比;与该假设一致但尚未被证实的是:确定性的带外强制执行相比带内检测,是更难被自适应攻击者攻破的目标。This work organizes out-of-band defenses as instances of classical integrity protection, reference monitoring, and least privilege, yielding a structured comparison of what they do and do not cover, consistent with, but not established, the hypothesis that deterministic out-of-band enforcement is a harder target for an adaptive attacker than in-band detection.
本文展示了用于 IT-Grundschutz(IT-GS)认证部分自动化的多 Agent 系统(MAS)架构结合混合检索增强生成(HybridRAG)的技术实现与实证评估,并为强化合规严谨性引入两项新的 MAS 架构技术贡献。This paper presents the technical implementation and empirical evaluation of a Multi-Agent System (MAS) architecture combined with Hybrid Retrieval Augmented Generation (HybridRAG) for the partial automation of IT-GS certification and introduces two novel technical contributions to the MAS architecture to enforce the compliance rigor.
结果表明,工具使用评测应从函数调用准确率转向不可靠工具环境下的任务完成度,并建议工具使用评测应从函数调用准确率转向不可靠工具环境下的任务完成度。(注:原文末句疑似重复)Results suggest that tool-use evaluation should move beyond function-call accuracy toward task completion under unreliable tool environments, and suggest that tool-use evaluation should move beyond function-call accuracy toward task completion under unreliable tool environments.