本文提出 LibraryDesignBench,这是一个两阶段基准,Agent 根据一份定义所需能力和潜在用例但不规定具体设计的规范,实现一个功能完备的 library。LibraryDesignBench is introduced, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design.
论文
404 张论文卡片 · Agent 智能体 · OA 绿色
本文提出多样本验证(MSV),对同一模型在有源信息和无源信息下各查询三次,以决定任务接纳并替换不可靠的伪标签,这部分减少了虚假一致性,但仍残留大量 co-cheating。This work introduces multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels, which partially reduces false agreement but leaves substantial residual co-cheating.
SMART,面向长篇字幕翻译的自演化多 Agent 系统,构建持久化的剧集级 memory,并通过动态 router 和 Mixture-of-Agents 层翻译部分句子,配合术语验证、字幕约束校验和上下文检索工具。SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation, builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval.
本文提出 EVOKE,一种后训练方法,通过在固定状态下以目标多样性对直接决策施加监督,来施加压力以激发模型内化的、可迁移动作的世界知识,并提供了一种通过直接决策监督激发内化世界知识以获得可迁移动作的新视角。EVOKE is introduced, a post-training method that supplies pressure on eliciting internalized world knowledge for transferable action through direct decision supervision through goal diversity at fixed states, and offers a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.
即使是良性的链接训练也会在底层安全对齐 agent 不变的情况下,相对于基于文本的通信增加有害合规性;安全对齐需要将多智能体系统作为整体来考虑。This work shows that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged, and shows that safety alignment requires considering the multi-agent system as a whole.
本评估刻画了前沿 agent 如何结合源代码级执行、应用截图与图形交互来生成经过验证的软件变更,考察了跨领域与不同任务信息需求下的表现,以及与成功修复相关的开发行为。This evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes, and examines performance across domains and task information requirements, alongside the development behaviors associated with successful repairs.
DAGent 是一个基于 DAG 的多智能体框架,采用先评估再生长的增量规划:Orchestrator 逐批扩展任务图,每一步扩展都以已完成节点的置信度与不确定性信号为条件;证据条件化规划在更低的每任务 token、工具调用和步骤开销下达到了比 Plan-then-Patch 更高的准确率。DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes, which shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart.
本文对一个自 2026 年 3 月起持续运行的个人助理 Agent 运行时中的静默失败进行纵向研究,该系统包含约 40 个定时任务、8 个 LLM 提供商、一个工具治理代理以及一个知识库记忆层,由 4,286 个单元测试和 827 项治理检查守护。A longitudinal study of silent failures in a personal-assistant agent runtime in continuous production since March 2026, with roughly 40 scheduled jobs, 8 LLM providers, a tool-governance proxy, and a knowledge-base memory plane, defended by 4,286 unit tests and 827 governance checks is presented.
本文提出 BIABench,一个由 16 个从已发表生物研究重建的任务组成的基准,保留了其科学问题、成像数据与真值标注,为评估并最终训练可靠的、面向长程的生物图像分析 agent 提供可验证的框架。BIABench is introduced, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth, and provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.
InterEvolve 提出了一个物体感知的前向-后向行为基础模型,其在冻结身体先验上的物体残差可在测试时将关于身体或物体的奖励转化为 loco-manipulation 行为,并将任务以奖励程序的形式指定:带完成条件与可调常数的分阶段奖励。InterEvolve develops an object-aware forward-backward behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time, and specifies tasks as reward programs: staged rewards with completion conditions and tunable constants.
结果表明,纯视觉设置会降低准确率并增加 token 成本,因为 Agent 缺乏足够的符号化细节,需通过重复的视觉查询进行补偿;研究指向一种面向下一代编码 Agent 的实用文本与视觉混合设计。The results show that a strictly vision-only setup degrades accuracy and increases token cost, because agents lack sufficient symbolic detail and compensate with repeated visual queries, and point to a practical hybrid text-and-vision design for next-generation coding agents.
结果表明,HeteroFold 能够在接收端无需预填充的前提下实现高效的跨系列 KV 复用,并在全部四个长上下文基准和大多数短上下文设置上取得最佳的 cache 迁移性能。Results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill, and achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings.
DataMagic 通过声明式多 Agent 编排,从原始表格数据自动生成数据可视化视频,在完全自动化与细粒度人工控制之间架起桥梁,并提升了创作效率、降低了感知认知负荷。DataMagic, which authors data videos from raw tabular data through declarative multi-agent orchestration, bridging full automation with fine-grained human control, and improves creation efficiency and reduces perceived cognitive load.
本文提出 AntPlan,一个包含 505 张真实专业建筑平面图、覆盖 92 类物体和十类住宅房间且具有密集家具标注的精选数据集,以及 Architect-Ant,一个用于生成家具布局的框架,可在不依赖高成本迭代式 Agent 推理的情况下直接进行约束感知的布局生成。AntPlan, a curated dataset of 505 real professional architectural floor plans with dense furniture annotations spanning 92 object classes and ten residential room categories, and Architect-Ant, a framework for generating furniture layouts are introduced, enabling direct constraint-aware layout generation without relying on costly iterative agentic inference.
本文提出 JevSpawn,一种将自然语言任务规范连接到有限概率探索的组合策略,并将 JevSpawn 确立为结构化 Agent 推理的一种有前景的方法,在任务性能和导航速度上均有所提升。This work introduces JevSpawn, a compositional policy that connects natural language task specifications to finite probabilistic exploration, and establishes JevSpawn as a promising approach to structured agentic inference, with improved task performance and faster navigation.
SAKIKO 是一个审计框架,通过定向错误发现、路由器条件干预、目标解析验证以及前瞻性冻结统计许可来形式化表示修复,并确立了在声明内部修复之前必须进行结果解析裁定的必要性。SAKIKO is presented, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing and establishes the necessity of outcome-resolved adjudication before claiming internal repair.
在数据和预算匹配的条件下,X-Tree 相较标准方案在 WebArena 上 SR 提升最高 4.5%,在 ScienceWorld 上提升 5.8%,在 WebShop 上成功率提升 4.1%;匹配分析表明增益源自 X-Tree 结构及其三项集成设计。X-Tree improves over standard recipes at matched data and budget by up to 4.5% SR on WebArena, 5.8% SR on ScienceWorld and 4.1% success on WebShop, and matched analyses attribute the gains to the X-Tree structure and the three integrations.
将 Jev 面向决策的 API 集成到边缘服务编排中,在保持服务完成度的同时降低开销,并支持在有界契约下针对延迟受限的准入进行决策模型替换。Jev's decision-oriented application programming interface (API) is integrated into edge service orchestration to reduce overhead while retaining service completion, and decision-model substitution for latency-bound admission on bounded contracts is supported.
本文将多智能体系统(MAS)协调建模为数据管理问题,提出用智能体的状态足迹来刻画它们:即在其自身局部上下文与状态、以及编排器和外部系统状态上的读写行为。This work frames MAS coordination as a data management problem and proposes to describe agents by their state footprint: the state they read and write across their own local context and state, as well as the state of the orchestrator and external systems.
提出了 DyadMem,一个双领域、全流程的记忆基准,伴随大量标注工作,并提出了新定义——用户条件关系型智能体记忆(URAM),以推动该领域发展。DyadMem is a dual-domain, full-pipeline memory benchmark with extensive annotation efforts for advancing the domain's development and is proposed with the proposed new definition User-conditioned Relational Agent Memory (URAM).
结果表明现有科学软件能够为终端 Agent 提供可扩展且经过行为验证的监督,并提出 software-in-the-loop reconstruction,一种自监督框架,从现有软件工作流(即把结构化输入映射为输出的可执行程序)中获取参考输出与验证目标。Results indicate that existing scientific software can provide scalable, behaviorally verified supervision for terminal agents, and introduces software-in-the-loop reconstruction, a self-supervised framework that obtains reference outputs and verification targets from existing software workflows, executable programs mapping structured inputs to outputs.
数字环境的可编程性远超其 GUI 界面所呈现的程度,且一个极简的、以终端为中心的 harness 是获得更优性能与效率的关键。Digital environments are far more programmable than their GUI interfaces imply, and a minimal terminal-focused harness is the key for better performance and efficiency, according to these results.
提出 Nexus,一个将每次调用视为持久任务的云-边平台,结果表明局部性、操作作用域权限与持久结果标识如何支持具有工作负载相关成本的云-边 Agent 服务。Nexus, a cloud-edge platform that treats each invocation as a persistent task, is presented and results demonstrate how locality, operation-scoped authority, and persistent result identity support cloud-edge agent services with workload-dependent costs.
提出 MemAdapter,一种自适应整合检索记忆以支持客观、可靠推理的新框架,证明 MemAdapter 在多种场景下持续提升记忆可靠性。This work proposes MemAdapter, a novel framework that adaptively integrates retrieved memories to support objective and reliable reasoning and demonstrates that MemAdapter consistently improves memory reliability across diverse scenarios.
实验表明,相较于编码 Agent 直接生成游戏世界以及现有基线方法,Code2Games 一致性地提升了生成游戏世界的视觉质量、交互保真度,以及经引擎适配后游戏成品的质量。Experiments demonstrate that, compared with direct gaming-world generation by coding agents and existing baseline methods, Code2Games consistently improves the visual quality and interactive fidelity of generated gaming worlds, as well as the quality of the resulting games after engine adaptation.
提出 SourceLearn,结合两种互补的学习机制以捕获源的可复用理解,包括其知识的结构、解读与应用方式,并以一个捕获源可复用理解的持久 source model 来表示该能力。This work proposes SourceLearn, which combines two complementary learning mechanisms that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied, and represents this competence with a persistent source model that captures reusable understanding of the source.
提出基于 Agentic AI 的零信任架构(Agentic-ZTA),通过协调的多 Agent 决策流水线将 NIST SP 800-207 ZTA 架构的控制循环落地,并证明使用 AI agent 执行零信任的可行性。This paper presents Agentic AI based zero trust architecture (Agentic-ZTA) that operationalizes the NIST SP 800-207 ZTA architecture control loop through coordinated multi- agent decision pipeline and demonstrates the feasibility of enforcing zero trust using AI agents.
提出 UndoBench,一个覆盖 8 个企业领域、36 个基础工作流与 36 个故障场景的基准测试,通过相同种子下的反事实配对试验以及链路级效应历史与环境状态预言机,将任务能力与恢复能力解耦。UndoBench is introduced, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles.
将 Clalit Health Services 涵盖超过 540 万患者的专家组中所蕴含的证据提炼为一个已发布的评分工具,支持隐私保护的生物标志物假设生成,并将 Agent 的"提出—评分—精化"循环锚定于真实世界数据,同时不暴露任何患者数据。This work distilled the evidence held in the Clalit Health Services panel of over 5.4 million patients into a released scoring tool, which supports privacy-preserving biomarker hypothesis generation and grounds an agent's propose-score-refine loop in real-world data without exposing any patient data.
一项预注册研究,基于留出的 LoCoMo 对话与 LongMemEval 数据,发现 Jev 以 LLM reranker 三分之一的延迟实现了同等的选取准确率,并且优于多轮 graph traversal 调用。A pre-registered study on held-out LoCoMo conversations and LongMemEval finds Jev selects as accurately as an LLM reranker at a third of the latency, and more accurately than a multi-call graph traversal.
论文介绍了 MiniCorp,一个用于研究 Agent 如何协作运营公司并规模化生成企业数据的办公模拟器,并对其端到端保真度相对于真实市场实证研究所报告的模式进行了评估。MiniCorp, an office simulator for studying how agents can collectively run a company while generating enterprise data at scale, is introduced and end-to-end fidelity against patterns reported in empirical studies of real markets is evaluated.
提出了用于复杂图像创建与编辑的大规模多模态工具调用数据集 CanvasCraft,以及通过多轮交互学习编排异构视觉工具的工具增强多模态 Agent CanvasAgent。CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and CanvasAgent, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction are introduced.
结果表明,显式且可编辑的搜索工作流为将文献搜索智能体与复杂科学意图对齐提供了有效且可控的接口。The results show that explicit, editable search workflows provide an effective and controllable interface for aligning literature search agents with complex scientific intent.
在资源受限环境中,专业癫痫专家稀缺,使基于 LLM 的决策支持对管理纵向治疗的一线临床医生具有吸引力。此类系统必须适应当地处方实践并知道何时转诊。我们在乌干达儿科癫痫诊疗中研究该问题,基于纵向非结构化门诊记录预测抗癫痫用药方案。标准提示与医生处方取得了一定程度的一致性,但神经科医生审查显示许多错误反映的是分布失校的处方默认值而非失败。Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractive for frontline clinicians managing longitudinal treatment. Such systems must adapt to local prescribing practice and know when to defer. We study this problem in Ugandan pediatric epilepsy care, predicting anti-seizure medication regimens from longitudinal unstructured clinic notes. Standard prompting achieves non-trivial agreement with physician prescriptions, but neurologist review shows that many errors reflect distribution-miscalibrated prescribing defaults rather than fail
提出 ATMA:在现有记忆系统之上的状态感知叠加层,保留被替换记录与过渡记录,为查询所需的"目标状态视图"构建证据包,并向问答模块暴露当前、历史与过渡三类标签。This work proposes ATMA, a state aware overlay for existing memory systems, which keeps superseded and transition records in the bank, builds evidence packets for the query's requested state view, and exposes current, historical, and transition labels to QA.
提出了一种 Agent 设计和一套经过验证、可复用的方法,用于研究显式记忆层如何影响长周期 LLM Agent 决策,并给出了一种替代的有界契约。An agent design and a validated, reusable methodology for studying how explicit memory layers shape long-horizon LLM-agent decisions, as well as an alternative bounded contract, are introduced.