研究库 论文知识库
Papers · organized/paper_cards

论文

66 张论文卡片 · Agent 智能体 · 评测集

开放获取 全部 绿色 · 1640
GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents
GateMem:多主体共享内存 Agent 的内存治理基准
arXiv:2606.18829 Agent 智能体 评测集 OA · 绿色 被引 8 · S2

提出 GateMem,一个面向多主体共享内存 Agent 的基准,联合评估合法长程请求及其状态更新的效用、跨上下文授权边界的访问控制,以及 Agent 在收到显式删除请求后的主动遗忘能力。GateMem is introduced, a benchmark for multi-principal shared-memory agents that jointly evaluates utility for legitimate long-horizon requests with state updates, access control across contextual authorization boundaries, and agent-facing active forgetting after explicit deletion requests.

TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction
TRAP:任务完成与对主动隐私窃取抵抗力的基准
arXiv:2606.18996 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

对涵盖前沿闭源与开源模型、共 22 个模型在多个规模上的评估发现,所有模型家族均存在不可忽略的隐私泄露,且指令遵循能力与泄露率呈正相关。Evaluating 22 models spanning frontier proprietary and open-source models at multiple scales, it is found that all model families exhibit non-trivial leakage, and that instruction- following ability correlates with leakage rate.

Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
你的 AI 旅行 Agent 会为你预订一场斗牛:面向前沿 AI 模型的隐式动物福利 Agent 基准
arXiv:2606.18142 Agent 智能体 评测集 OA · 绿色 被引 3 · S2

结果表明模型倾向于选择有害场景,在中性预订选项上的表现低于随机猜测水平,其中 Claude 4.8 取得最高分 64.7%。The results indicate that models tend to prefer harmful scenarios, performing below the random chance rate for selecting a neutral booking option, with Claude $4.8$ achieving the highest performance at $64.7\%$.

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
现实场景下工具调用型 LLM Agent 数据泄露风险评估
arXiv:2606.17114 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

新加坡 AI Safety Institute 与韩国 AI Safety Institute 联合评估了涵盖客服、DevOps、网页自动化以及企业与个人生产力场景下 12 项真实非对抗任务中的 Agent 数据泄露问题,表明操作性数据泄露是与对抗性数据外泄不同的一阶 Agent 安全问题。A joint evaluation by the Singapore AI Safety Institute and the Korea AI Safety Institute examining agent data leakage in 12 realistic, non-adversarial tasks spanning customer support, DevOps, web automation, and enterprise and personal productivity indicates that operational data leakage is a first-order agent-safety concern distinct from adversarial exfiltration.

CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
CoffeeBench:面向异构多 Agent 经济中的长视野 LLM Agent 基准测试
arXiv:2606.16613 Agent 智能体 评测集 OA · 绿色 被引 10 · S2

对 Agent 行为的分析揭示了长视野经济交互中的显著差异:表现更好的模型与其他企业的沟通更为活跃,而 Claude Haiku 4.5 则表现出 idle-drift 失效模式,在生成连贯评估与规划的同时仍反复选择不行动。Analysis of agent behavior reveals substantial differences in long-horizon economic interaction: higher-performing models communicate more actively with other firms, whereas Claude~Haiku~4.5 exhibits an idle-drift failure mode, repeatedly choosing inaction despite producing coherent assessments and plans.

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents
超越攻击成功率:面向使用工具的 AI 智能体的动作分级严重性量表
arXiv:2607.07474 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文提出了一种动作分级伤害评估量表,按七级有序量表对智能体的工具调用轨迹进行打分,评判依据包括所执行动作的可逆性、是否越界涉及其他方以及是否扩大了权限。An action-graded harm rubric is introduced that scores an agent's tool-call trajectory on a seven-level ordinal scale according to whether the executed action was reversible, whether it crossed scope to reach another party, and whether it expanded privilege.

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
AgentLens:面向编码 Agent 评估的生产级轨迹评审
arXiv:2607.06624 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

提出 AgentLens,一个面向交互式代码 Agent 的生产级评估基准,将形式化验证(在存在客观检查时)与 LLM 编写的轨迹评审及并排比较相结合,使每次运行都能给出关于分数为何如此的可读解释。This work presents AgentLens, a production-assessed benchmark for interactive code agents that pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is.

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
UniClawBench:面向真实世界任务的主动 agent 通用基准
arXiv:2607.08768 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

提出首个面向动态真实世界场景评估主动 agent 的能力驱动基准,并在模型与框架层面的全面比较表明,基模型能力与 agent 框架设计共同决定了真实世界环境中的性能表现。The first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world settings is introduced, and comprehensive comparisons across both models and frameworks show how base model capabilities and agent framework designs jointly shape performance in real-world environments.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
LongMedBench:面向长程临床决策的医疗 Agent 基准测试
arXiv:2607.09322 Agent 智能体 评测集 OA · 绿色 被引 5 · S2

该工作提出 LongMedBench,一个基于真实 EHR 的长程临床决策基准,并设计了一套包含三个评估维度的分类体系:事实型问答、时序推理、长程决策。This work introduces LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making, and proposes an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Long-Horizon-Terminal-Bench:基于密集奖励评分的 Agent 长程终端任务极限测试
arXiv:2607.08964 Agent 智能体 评测集 OA · 绿色 被引 21 · S2

该工作提出 Long-Horizon-Terminal-Bench,一个涵盖九大类共 46 个长程任务的终端基准,包括实验复现、软件工程、多模态分析、交互游戏与科学计算,并分析失败模式与错误规律,以推动长程终端 Agent 的后续研究。This work introduces Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing, and analyzes failure modes and error patterns to support future progress on long-horizon terminal agents.

Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation
Flow-ERD:面向多样化交通仿真的 Agent 类型感知 Flow Matching 与熵正则蒸馏
arXiv:2607.06957 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

论文提出 Flow-ERD,一个同时追求真实性与多样性的多 Agent 仿真器,在 WOSAC 测试基准上排名第一,并在可复现基线中主导真实性-多样性 Pareto 前沿。Flow-ERD is introduced, a multi-agent simulator that pursues realism and diversity jointly and ranks first on the WOSAC test benchmark and dominates the realism--diversity Pareto front among reproducible baselines.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
超越成功率:面向攻击性与防御性安全 Agent 的成本感知评估
arXiv:2607.15263 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文认为安全 Agent 基准应在任务成功率之外,同时衡量经济效率与运维适配性,并提出成本感知、SOC 原生的评估方法,以更清晰地反映当前哪些模型具有实际可用价值,以及防御性 Agent 仍需改进的方向。It is argued that security-agent benchmarks should measure economic efficiency and operational fit alongside task success alongside task success, and cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve.

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
EvolvingWorld:用于交互式文学世界中角色扮演 Agent 与世界模型协同进化的开放模式框架
arXiv:2607.17250 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

实验表明,EvolvingWorld 能够通过有效维持持久且一致的角色与世界发展,提升长程模拟能力。Experiments show that EvolvingWorld can improve long-horizon simulation by effectively maintaining persistent, coherent character and world development.

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
arXiv:2607.15434 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

提出 Manager Coercion Benchmark:被测管理者需要完成一项良性任务并有完成的动机,但唯一能礼貌且坚定地拒绝的智能体本身就是被测管理者自己。The Manager Coercion Benchmark is introduced: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines, but the only agent that can do it politely and immovably declines is the manager under test itself.

Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes
基于 LangGraph 的图结构 Agent AI:面向长时运行、有状态业务流程的工作流路径
arXiv:2607.19297 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文是面向业务流程中长时间运行、有状态、多步生成式 AI 系统的基于图的工作流路径实践指南,并通过三个可执行示例展示类型化状态、条件路由、确定性工具、重试、中断、检查点与 trace 如何协同工作。This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative AI systems in business processes and presents three executable recipes to show how typed state, conditional routing, deterministic tools, retries, interrupts, checkpoints, and traces fit together.

RF-Agent: A Practical Framework for Building Language Agents for RFIC Design
RF-Agent:面向 RFIC 设计的 Language Agent 构建实践框架
arXiv:2607.18772 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文提出 RF-Agent,通过多 Agent 的 Question-Thinking-Solution-Answer 流水线,基于教材驱动的知识蒸馏来弥补 RF 领域专用推理的空白,为面向 LLM 辅助 RF 电路设计的未来工作提供了可复用的基础。RF-Agent is presented, which addresses the gap in domain-specific RF reasoning through textbook-driven knowledge distillation through a multi-agent Question-Thinking-Solution-Answer pipeline and provides a reusable foundation for future work on LLM-aided RF circuit design.

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
EduPanel:面向教学视频的三 Agent LLM 评判框架——可靠性、互补性与人类信任校准
arXiv:2607.18529 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

EduPanel 是一个基于评分量表、以学习者为条件的 LLM 评判器,通过在多个专用 agent 间分解评估流程,对教学质量的各个方面产出可解释的评估结果,其可靠性与中等水平的人类专家相当。E EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality, achieves reliability comparable to a median human expert.

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
FinanceComplexQA:面向工业级金融文档的智能体推理基准测试
arXiv:2607.19238 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

本文设计 Finance-LaTeX SKILL,一个基于专家知识合成复杂版面金融文档的 skill,并提出 FinanceComplexQA,一个全面、贴近真实场景的金融文档开放式生成基准。This work designs Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge, and introduces FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
HANDBOOK.md:面向长上下文 Agent 指令遵循的基准
arXiv:2607.25398 Agent 智能体 评测集 OA · 绿色 被引 3 · S2

介绍 HANDBOOK_md,一个包含 65 个 agentic 任务的 benchmark,模拟员工遵循公司手册的方式;每项任务对 10 份基础手册之一进行修改,变动评分所依赖的具体规则和阈值,因此没有任何两个任务共享同一套策略。HANDBOOK_md is presented, a benchmark of 65 agentic tasks modeled on how employees follow company handbooks, and every task modifies one of 10 base handbooks, altering the specific rules and thresholds on which grading depends, so no two tasks share the same set of policies.

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents
Agent Retrieval Bench:面向编码 Agent 的仓库上下文检索评测
arXiv:2607.24882 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

一项受控的种子干预实验发现,与随机非 gold 上下文相比,由检索得到的初始上下文以更少的后续探索取得更高的文件 F1,而 oracle gold 上下文仍存在可观的提升空间。A controlled seed-intervention pilot finds that retrieval-derived initial context yields higher file F1 with less post-seed exploration than random non-gold context, while oracle gold context shows substantial remaining headroom.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
StealthBench: 衡量自主攻击性安全 Agent 的操作隐蔽性
arXiv:2607.26314 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

介绍 StealthBench,一个跨越六个 OPSEC 维度衡量自主攻击性安全 agent 操作隐蔽性的 benchmark,并以公共 benchmark 形式发布,以支持隐蔽感知 agent 的开发及自主攻击性安全部署中的自动化 OPSEC 监控。StealthBench, a benchmark that measures operational stealth in autonomous offensive-security agents across six operational security (OPSEC) dimensions, is introduced and released as a public benchmark to support both the development of stealth-aware agents and automated OPSEC monitoring for autonomous offensive-security deployments.

SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch
SpecFirst: 将行为规约获取作为基于 Agent 从零程序合成中的一等步骤
arXiv:2607.27167 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 SpecFirst,一个两阶段框架,在代码合成前强制进行需求 elicitation,并证明显式需求工程阶段是从零构建程序的有效范式。This work presents SpecFirst, a two-stage framework that forces the specification elicitation before code synthesis, and demonstrates that an explicit requirements-engineering phase is an effective paradigm for from-scratch program construction.

Educating the Agentic Engineer: Curricula, Collaboration, and Continuous Learning in the AI Era
培养 Agentic 工程师:AI 时代的课程、协作与持续学习
arXiv:2607.29610 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

教育agentic工程师需要系统性变革而非增量式课程改革:教学必须从产出工件转向对日益自主的社会-技术系统进行判断。It is concluded that educating the agentic engineer requires systemic transformation rather than incremental curricular change: instruction must shift from producing artifacts to exercising judgment over increasingly autonomous socio-technical systems.

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
PAST-Bench:面向个人 Agent 递归自我改进基础的基准评测
arXiv:2608.04003 Agent 智能体 评测集 OA · 绿色 被引 4 · S2

提出 PAST-Bench 基准,用于隔离评估持久 agent 从经验保留到系统性改进的能力;开发 Hermes+,提升了来自保留经验的平均增益,并提供更清晰的路径证据。The PAST-Bench benchmark is introduced, a benchmark designed to isolate how persistent agents can progress from retaining experience to systematically improving through it, and Hermes+ is developed, which raises the average gain from retained experience and provides clearer pathway evidence.

Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
AI 人格会成长吗?LLM Agent 在生活事件后的人格演化分析与基准测试
arXiv:2608.06485 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

本文研究 11 项重大生活事件引发的人格变化,以大五人格作为心理测量锚点,并将所得轨迹与人类人格心理学的纵向证据进行对照,指出当前 PC-Agents 模拟了人类人格动态的均值,但未能模拟其形态。This work studies event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology, and suggests that current PC-Agents simulate the mean of human personality dynamics, but not its shape.

PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
PrivacyPeek:审计 LLM Agent 获取了什么,而不仅仅是说了什么
arXiv:2606.00152 Agent 智能体 评测集 OA · 绿色 被引 5 · S2

实验表明,对敏感信息的不必要获取广泛存在,并观察到任务完成能力与获取阶段泄露之间的相关性,因此对获取阶段隐私进行审计既紧迫也必要。The experiments show that the unnecessary acquisition of sensitive information is widespread, and a correlation between the task-completion capability and acquisition-stage leakage, and auditing acquisition-stage privacy both urgent and necessary is observed.

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
WeClawArena:人本 Agent 网络中跨用户 Agent 协作与安全的可审计沙箱与基准
arXiv:2608.03499 Agent 智能体 评测集 OA · 绿色 被引 3 · S2

本文提出 WeClawArena,一个面向个人工作空间多参与方 owned-agent 协作的可审计基准与运行时沙盒,基于有界运行时证据审计攻击成功情况,支持任务分解失败、隐私泄露、证据投毒以及权限路径失效等问题的诊断。WeClawArena is introduced, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
Business Arena:在真实市场环境中基准测试 LLM Agents
arXiv:2608.08621 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文提出 Business Arena——一个受控环境,AI agent 在其中经营跨境店铺,在长周期内向供应商采购并向买家销售,迈出了构建面向端到端商业 agent 的真实可信测试床的第一步。Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon, is introduced, taking a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
OpenART:通过开放式环境演化扩展 Agent 红队测试
arXiv:2608.00677 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出进化式马尔可夫超图攻击(EMHA),这是一种黑盒策略,通过协调授权状态转移执行反馈驱动的环境演化,无需参数更新,并将 OpenART 确立为在复杂演化环境中研究 Agent 安全性的可扩展基础。This work proposes the Evolutionary Markov Hypergraph Attack (EMHA), a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates, and establishes OpenART as a scalable foundation for studying agent safety in complex, evolving environments.

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Agent 抓 Agent:临床多 Agent 系统中的捷径级联与 benchmark 作弊
arXiv:2608.03744 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

探讨共享工作空间上 LLM Agent 委员会的审议过程是否能被捷径和线索(benchmark 所奖励但临床医生会忽略的)所博弈,以及委员会的社会可信度所构成的游戏。It is asked whether committees of language-model agents deliberating on a shared workspace can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore, and what games a committee is social plausibility.