提出 PAST-Bench 基准,用于隔离评估持久 agent 从经验保留到系统性改进的能力;开发 Hermes+,提升了来自保留经验的平均增益,并提供更清晰的路径证据。The PAST-Bench benchmark is introduced, a benchmark designed to isolate how persistent agents can progress from retaining experience to systematically improving through it, and Hermes+ is developed, which raises the average gain from retained experience and provides clearer pathway evidence.
论文
43 张论文卡片 · Agent 智能体 · 评测集
本文研究 11 项重大生活事件引发的人格变化,以大五人格作为心理测量锚点,并将所得轨迹与人类人格心理学的纵向证据进行对照,指出当前 PC-Agents 模拟了人类人格动态的均值,但未能模拟其形态。This work studies event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology, and suggests that current PC-Agents simulate the mean of human personality dynamics, but not its shape.
实验表明,对敏感信息的不必要获取广泛存在,并观察到任务完成能力与获取阶段泄露之间的相关性,因此对获取阶段隐私进行审计既紧迫也必要。The experiments show that the unnecessary acquisition of sensitive information is widespread, and a correlation between the task-completion capability and acquisition-stage leakage, and auditing acquisition-stage privacy both urgent and necessary is observed.
本文提出 WeClawArena,一个面向个人工作空间多参与方 owned-agent 协作的可审计基准与运行时沙盒,基于有界运行时证据审计攻击成功情况,支持任务分解失败、隐私泄露、证据投毒以及权限路径失效等问题的诊断。WeClawArena is introduced, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.
本文提出 Business Arena——一个受控环境,AI agent 在其中经营跨境店铺,在长周期内向供应商采购并向买家销售,迈出了构建面向端到端商业 agent 的真实可信测试床的第一步。Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon, is introduced, taking a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.
本文提出进化式马尔可夫超图攻击(EMHA),这是一种黑盒策略,通过协调授权状态转移执行反馈驱动的环境演化,无需参数更新,并将 OpenART 确立为在复杂演化环境中研究 Agent 安全性的可扩展基础。This work proposes the Evolutionary Markov Hypergraph Attack (EMHA), a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates, and establishes OpenART as a scalable foundation for studying agent safety in complex, evolving environments.
探讨共享工作空间上 LLM Agent 委员会的审议过程是否能被捷径和线索(benchmark 所奖励但临床医生会忽略的)所博弈,以及委员会的社会可信度所构成的游戏。It is asked whether committees of language-model agents deliberating on a shared workspace can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore, and what games a committee is social plausibility.