D-JEPA 是一种决策对齐的潜在世界模型,从已执行结果中学习候选未来之间的决策相关关系,将决策相关关系结构确立为预测世界建模与有效控制之间的直接桥梁。D-JEPA is introduced, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes, establishing decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.
论文
284 张论文卡片 · Agent 智能体 · 方法
Just-In-Time Agent Memory (JAM) 是一个可训练框架,在运行时进行查询条件下的上下文构建,其任务性能优于 AOT 式记忆系统,同时比先前的可训练 Agent 记忆方法显著更高效。Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime, is proposed, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches.
Stashbird 是一种 Agent 记忆系统,通过显式来源溯源将源事件链接到派生记忆状态,在 LongMemEval-S 和 GroupMemBench 上准确率高于 Hindsight,在 EverMemBench 上与之相当。Stashbird is presented, an agent memory system that links source episodes to derived memory state through explicit provenance and achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.
本文提出 WideSWE,用于在跨仓库任务上评估编码 Agent,并在相同 prompt 下将其与联合执行进行对比,以考察逐仓库工作是否能缓解相关困难。This work introduces WideSWE to evaluate coding agents on cross-repository tasks, and compares it with joint execution under identical prompts to examine whether working on one repository at a time can alleviate difficulties.
提出 GAGAR,一个面向代码 agent RL 的质量感知信用再分配框架;结果表明,将基于测试的验证与成组的 agentic 评分相结合,可提升代码 agent RL 的质量与稳定性。GAGAR, a framework for quality-aware credit redistribution in code agent RL, is introduced and results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.
介绍 Relic,它将反复出现的协作失败转化为组织拥有的、可执行的协议,使协作经验成为超越其创建者仍持续有用的组织级持久状态。The introduction of Relic, which turns recurring collaboration failures into organization-owned, executable protocols, and how collaboration experience can become persistent organizational state that remains useful beyond the members who created it are shown.
提出 Context Language Models,能原生管理自身上下文,并可通过经标准 skill 优化循环演化的自然语言指令进行引导,在上下文管理任务上将未见数据的准确率最高提升 35.9 分,同时降低计算开销。Context Language Models are introduced, language models that natively manage their own context and can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute.
提出 CorpusMap,一个以语料中重复出现的实体为核心的导航层;这些实体可从文档自身识别,并能在不同来源间把单个文档与众多其他文档相连,表明实体可作为大型文档集合导航的有效锚点。CorpusMap is introduced, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources, suggesting that entities serve as effective anchors for navigating large document collections.
在所测试的每一对基座与指令微调模型中,指令微调都强化了模型对保留标记的偏好,且该差距在该通道上持续存在。In every base and instruction-tuned pair the authors test, instruction tuning strengthens the model's preference for reserved markers, and the gap persists on that channel.
KUPAS MASTER 是一个围绕九层认知语料构建的经验工程平台,将异构的工作记录与从业者访谈转化为 agent 可用的可追溯、可复用经验语料,提供了从个体隐性经验到组织知识与 agent 能力的可行路径。KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction, turns heterogeneous work records and practitioner interviews into traceable, reusable experience corpora for agents, and provides a practical path from individual tacit experience to organizational knowledge and agent capabilities.
本文提出一种异步 LLM 框架,允许用户(或 Agent 自身)定义具有重叠 memory 状态的推理协程,并展示 Qwen 3.x 模型无需任务专属训练即可在流式视频理解、电子游戏和监控中实现异步运行。This work develops an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states and showcases that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.
本文提出 Org-Agent,一种以约束为中心的统一推理框架,将任务执行组织为三个阶段,并将任务分解为原子子任务,构建编码其依赖关系的任务依赖图。Org-Agent is introduced, a unified constraint-centric reasoning framework that organizes task execution in three stages and decomposes a task into atomic subtasks and constructs a task dependency graph whose edges encode the dependencies among them.
本文提出多样本验证(MSV),对同一模型在有源信息和无源信息下各查询三次,以决定任务接纳并替换不可靠的伪标签,这部分减少了虚假一致性,但仍残留大量 co-cheating。This work introduces multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels, which partially reduces false agreement but leaves substantial residual co-cheating.
SMART,面向长篇字幕翻译的自演化多 Agent 系统,构建持久化的剧集级 memory,并通过动态 router 和 Mixture-of-Agents 层翻译部分句子,配合术语验证、字幕约束校验和上下文检索工具。SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation, builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval.
即使是良性的链接训练也会在底层安全对齐 agent 不变的情况下,相对于基于文本的通信增加有害合规性;安全对齐需要将多智能体系统作为整体来考虑。This work shows that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged, and shows that safety alignment requires considering the multi-agent system as a whole.
DAGent 是一个基于 DAG 的多智能体框架,采用先评估再生长的增量规划:Orchestrator 逐批扩展任务图,每一步扩展都以已完成节点的置信度与不确定性信号为条件;证据条件化规划在更低的每任务 token、工具调用和步骤开销下达到了比 Plan-then-Patch 更高的准确率。DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes, which shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart.
本文对一个自 2026 年 3 月起持续运行的个人助理 Agent 运行时中的静默失败进行纵向研究,该系统包含约 40 个定时任务、8 个 LLM 提供商、一个工具治理代理以及一个知识库记忆层,由 4,286 个单元测试和 827 项治理检查守护。A longitudinal study of silent failures in a personal-assistant agent runtime in continuous production since March 2026, with roughly 40 scheduled jobs, 8 LLM providers, a tool-governance proxy, and a knowledge-base memory plane, defended by 4,286 unit tests and 827 governance checks is presented.
InterEvolve 提出了一个物体感知的前向-后向行为基础模型,其在冻结身体先验上的物体残差可在测试时将关于身体或物体的奖励转化为 loco-manipulation 行为,并将任务以奖励程序的形式指定:带完成条件与可调常数的分阶段奖励。InterEvolve develops an object-aware forward-backward behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time, and specifies tasks as reward programs: staged rewards with completion conditions and tunable constants.
结果表明,纯视觉设置会降低准确率并增加 token 成本,因为 Agent 缺乏足够的符号化细节,需通过重复的视觉查询进行补偿;研究指向一种面向下一代编码 Agent 的实用文本与视觉混合设计。The results show that a strictly vision-only setup degrades accuracy and increases token cost, because agents lack sufficient symbolic detail and compensate with repeated visual queries, and point to a practical hybrid text-and-vision design for next-generation coding agents.
结果表明,HeteroFold 能够在接收端无需预填充的前提下实现高效的跨系列 KV 复用,并在全部四个长上下文基准和大多数短上下文设置上取得最佳的 cache 迁移性能。Results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill, and achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings.
DataMagic 通过声明式多 Agent 编排,从原始表格数据自动生成数据可视化视频,在完全自动化与细粒度人工控制之间架起桥梁,并提升了创作效率、降低了感知认知负荷。DataMagic, which authors data videos from raw tabular data through declarative multi-agent orchestration, bridging full automation with fine-grained human control, and improves creation efficiency and reduces perceived cognitive load.
本文提出 AntPlan,一个包含 505 张真实专业建筑平面图、覆盖 92 类物体和十类住宅房间且具有密集家具标注的精选数据集,以及 Architect-Ant,一个用于生成家具布局的框架,可在不依赖高成本迭代式 Agent 推理的情况下直接进行约束感知的布局生成。AntPlan, a curated dataset of 505 real professional architectural floor plans with dense furniture annotations spanning 92 object classes and ten residential room categories, and Architect-Ant, a framework for generating furniture layouts are introduced, enabling direct constraint-aware layout generation without relying on costly iterative agentic inference.
SAKIKO 是一个审计框架,通过定向错误发现、路由器条件干预、目标解析验证以及前瞻性冻结统计许可来形式化表示修复,并确立了在声明内部修复之前必须进行结果解析裁定的必要性。SAKIKO is presented, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing and establishes the necessity of outcome-resolved adjudication before claiming internal repair.
在数据和预算匹配的条件下,X-Tree 相较标准方案在 WebArena 上 SR 提升最高 4.5%,在 ScienceWorld 上提升 5.8%,在 WebShop 上成功率提升 4.1%;匹配分析表明增益源自 X-Tree 结构及其三项集成设计。X-Tree improves over standard recipes at matched data and budget by up to 4.5% SR on WebArena, 5.8% SR on ScienceWorld and 4.1% success on WebShop, and matched analyses attribute the gains to the X-Tree structure and the three integrations.
数字环境的可编程性远超其 GUI 界面所呈现的程度,且一个极简的、以终端为中心的 harness 是获得更优性能与效率的关键。Digital environments are far more programmable than their GUI interfaces imply, and a minimal terminal-focused harness is the key for better performance and efficiency, according to these results.
提出 Nexus,一个将每次调用视为持久任务的云-边平台,结果表明局部性、操作作用域权限与持久结果标识如何支持具有工作负载相关成本的云-边 Agent 服务。Nexus, a cloud-edge platform that treats each invocation as a persistent task, is presented and results demonstrate how locality, operation-scoped authority, and persistent result identity support cloud-edge agent services with workload-dependent costs.
提出 MemAdapter,一种自适应整合检索记忆以支持客观、可靠推理的新框架,证明 MemAdapter 在多种场景下持续提升记忆可靠性。This work proposes MemAdapter, a novel framework that adaptively integrates retrieved memories to support objective and reliable reasoning and demonstrates that MemAdapter consistently improves memory reliability across diverse scenarios.
实验表明,相较于编码 Agent 直接生成游戏世界以及现有基线方法,Code2Games 一致性地提升了生成游戏世界的视觉质量、交互保真度,以及经引擎适配后游戏成品的质量。Experiments demonstrate that, compared with direct gaming-world generation by coding agents and existing baseline methods, Code2Games consistently improves the visual quality and interactive fidelity of generated gaming worlds, as well as the quality of the resulting games after engine adaptation.
提出 SourceLearn,结合两种互补的学习机制以捕获源的可复用理解,包括其知识的结构、解读与应用方式,并以一个捕获源可复用理解的持久 source model 来表示该能力。This work proposes SourceLearn, which combines two complementary learning mechanisms that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied, and represents this competence with a persistent source model that captures reusable understanding of the source.
提出基于 Agentic AI 的零信任架构(Agentic-ZTA),通过协调的多 Agent 决策流水线将 NIST SP 800-207 ZTA 架构的控制循环落地,并证明使用 AI agent 执行零信任的可行性。This paper presents Agentic AI based zero trust architecture (Agentic-ZTA) that operationalizes the NIST SP 800-207 ZTA architecture control loop through coordinated multi- agent decision pipeline and demonstrates the feasibility of enforcing zero trust using AI agents.
将 Clalit Health Services 涵盖超过 540 万患者的专家组中所蕴含的证据提炼为一个已发布的评分工具,支持隐私保护的生物标志物假设生成,并将 Agent 的"提出—评分—精化"循环锚定于真实世界数据,同时不暴露任何患者数据。This work distilled the evidence held in the Clalit Health Services panel of over 5.4 million patients into a released scoring tool, which supports privacy-preserving biomarker hypothesis generation and grounds an agent's propose-score-refine loop in real-world data without exposing any patient data.
一项预注册研究,基于留出的 LoCoMo 对话与 LongMemEval 数据,发现 Jev 以 LLM reranker 三分之一的延迟实现了同等的选取准确率,并且优于多轮 graph traversal 调用。A pre-registered study on held-out LoCoMo conversations and LongMemEval finds Jev selects as accurately as an LLM reranker at a third of the latency, and more accurately than a multi-call graph traversal.
论文提出 SafeActBench,包含跨 6 个操作域与 5 种协议的 656 个用例,协议由静态动作判定和经调查的非动作逐步演进至单动作与多动作工作流,结果表明失败不仅源于信息缺失,也源于 Agent 在决策和执行动作时如何使用既有证据。SafeActBench is introduced, comprising 656 cases across six operational domains and five protocols that progress from static action judgment and investigated non-action to single- and multi-action workflows, showing that failures arise not only from missing information, but also from how agents use established evidence when deciding and executing actions.
本综述聚焦在部署阶段为 LLM 增强此类参数化记忆的方法:一个承载记忆的参数对象在推理时被插入到前向传播中,无论该对象是在部署前还是部署中获得。This survey focuses on methods that augment LLMs with such parametric memory at deployment: a memory-bearing parameter object is plugged into the forward pass during inference, whether it is acquired before or during deployment.
论文介绍了 MiniCorp,一个用于研究 Agent 如何协作运营公司并规模化生成企业数据的办公模拟器,并对其端到端保真度相对于真实市场实证研究所报告的模式进行了评估。MiniCorp, an office simulator for studying how agents can collectively run a company while generating enterprise data at scale, is introduced and end-to-end fidelity against patterns reported in empirical studies of real markets is evaluated.
本文提出 DAEDALUS,一种无需现有任务或 oracle 验证器、从自生成实践中引导可复用 agent 记忆的方法,并证明在较小探索预算下即可获得性能提升,其启发式策略同样能惠及其他模型家族的 agent。DAEDALUS is presented, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers, and it is shown that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families.