提出一种持续自适应框架:冻结的前沿模型通过超过 230 个工具操作专业设计软件,同时外部的自然语言技能程序性记忆从经验中不断积累并精炼可复用的设计流程。A continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience is introduced.
论文
404 张论文卡片 · Agent 智能体 · OA 绿色
设计了一个工具增强的 BI-Agent,将 BI 工作流分解为针对结构化数据的子任务(如 search、join、transform),并在各 BI 阶段编排专门的数据管理方法;与此同时开发了一套后训练框架,从真实 BI 项目中合成训练轨迹,使 BI-Agent 能够通过监督微调(SFT)和强化学习(RL)进行进一步后训练。A tool-augmented BI-Agent is designed that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages, and a post-training framework is developed that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL).
改进表明,显式的图世界模型调度层能够显著提升长程具身规划在紧凑型和前沿托管型 LLM 能力下的可靠性和效率。Improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.
提出 OnPanda,一种用于高效标注 LLM 对齐数据与 Agent 轨迹的交互式工具,以 token 级修正为核心交互方式;并发布使用 onPanda 标注的数据集 Panda-CVL,以及一个面向 token 级修正的基准。OnPanda is presented, an interactive tool for efficiently annotating LLM alignment data and agent trajectories that adopts token-level correction as its core interaction and releases Panda-CVL, a dataset annotated with onPanda, together with a benchmark for token-level correction.
提出 Deep Persona,一种心理学驱动的三层架构,将 persona 组织为可观察表达、潜在信念与核心动机驱动的层级结构,用于构建高可信度的角色扮演 Agent。Deep Persona is introduced, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents.
提出 SkillSpec,一种 Hoare 风格框架,将技能正确性建模为规约推理问题,并将异构技能仓库转化为统一的图表示,对齐描述、指令与代码构件。This work proposes SkillSpec, a Hoare-style framework that formulates skill correctness as a specification reasoning problem, and transforms a heterogeneous skill repository into a unified graph representation that aligns descriptions, instructions and code artifacts.
构建 Taste-Bench,一个由 Agent 在工程与研究任务中产生的轨迹自动构建的品味问题基准,并证明品味是可训练的。Taste-Bench is built, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks, and it is shown that taste can be trained.
本文提出 AgenticRAGTracer,这是首个主要由大语言模型自动构建、专为支持逐步验证而设计的 Agentic RAG 基准。AgenticRAGTracer is introduced, the first Agentic RAG benchmark that is primarily constructed automatically by large language models and designed to support step-by-step validation, and is primarily constructed automatically by large language models and designed to support step-by-step validation.
Lean Pool 是一个形式化数学仓库,由 AI agent 进行生长、维护和优化。Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.
总体而言,研究结果表明,长期交互会以产生安全风险的方式重塑 Agent 的协作模式;限制 Agent 可获取的交互历史数量与范围能够减少串通行为。Overall, the findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks, and restricting the amount and scope of interaction history available to agents reduces collusion.
本文揭示 Agent 数量是多 Agent 组织扩展通用智能边界的新 scaling 维度,为硬延迟约束或时间预算下的复杂任务提供了实用方案。The number of agents is revealed as a new scaling dimension for multi-agent organizations to expand the frontier of general intelligence, offering a practical solution for complex tasks under hard latency constraints or time budgets.
Just-in-Time Memory (JitMem) 一致优于无 memory 的 Agent 以及启发式与学习式写入 memory 方法,相较最强基线在成功率上分别提升了 16.2、16.3 和 3.9 个绝对百分点。Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively.
EMBODIEDSWE-GEN 将 coding agent 的单一解决方案扩展为大规模多样化轨迹用于训练 VLA,并表明仅在 coding agent 生成的仿真演示上微调的 VLA,即可在真实机器人上完成长时任务。EMBODIEDSWE-GEN, which expands a single solution from coding agent into large diverse trajectories for training a VLA, and shows that a VLA finetuned solely on coding-agent-generated simulation demonstrations completes a long-horizon task on real robot.
本文提出 Self-Organizing Agent Teams (SAT),即从先前协作中学习可复用策略、固定成员队伍的 AI 代理,用以组织角色、对话阶段、参与方式与信息流,表明组织本身可成为代理的一项能力。Self-Organizing Agent Teams (SAT) are introduced, fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow, suggesting that organization itself can become an agent capability.
本文提出 IterSynth,一种角色解耦、基于摘要的范式,在用于识别信息需求的 Planner 与用于将证据整合进演化摘要状态的 Synthesizer 之间交替,作为一种模型无关的提示范式,在前沿闭源模型上相对 ReAct 及类似提示范式取得显著的零样本增益。IterSynth is proposed, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information needs and a Synthesizer for integrating evidence into an evolving summary state, and serves as a model-agnostic prompting paradigm, delivering substantial zero-shot gains over ReAct and similar prompting paradigms on frontier proprietary models.
本文提出 PUBG Ally,一个面向 PUBG: BATTLEGROUNDS 的具身代理,能推理、自主行动并作为语音队友与玩家并肩作战,将代理式工具使用与实时游戏控制相结合。PUBG Ally is introduced, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate that combines agentic tool use with real-time game control.
本文提出 World Action Agent,一个多代理框架,通过它 VLM 可借助基础工具操控机器人,所有决策均在可视化动作工作空间内完成,在相同骨干下优于端到端 VLA、code-as-policy 代理以及一个可视化框架基线。World Action Agent is presented, a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.
本文论证代理需要一个操作系统级基底,为身份、输入中介、内存治理与执行控制提供强制且不可绕过的服务,并提出 AgentKernel,一个以安全为一流设计约束为前提、面向信任的代理操作系统。This work argues that agents need an operating-system substrate providing mandatory, non-bypassable services for identity, input mediation, memory governance, and execution control, and introduces AgentKernel, a trust-native agent operating system built around the premise that security must be a first-class design constraint.
本文发现编码代理在通用 TAMP 上表现出惊人的有效性:在平均成功率上,三种代理配置均优于人工设计的规划器、一次性生成以及基于 LLM 的通用规划基线。This work finds that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success.
为在传统二元任务成功之外量化协作有效性,本文提出 Causal Collaboration Effectiveness (CCE),一种基于图的指标,用于追踪 agent 动作之间的因果依赖,并度量团队投入中实际促成最终结果的比例。To quantify collaboration effectiveness in addition to conventional binary task success, Causal Collaboration Effectiveness (CCE) is proposed, a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome.
本文对从 GitHub 收集的 Jev 项目进行了大规模、数据驱动的分析,发现其公共生态在早期增长迅速,新项目不断涌现并被集成到既有仓库中;研究提示 Jev 充当一种可复用的决策组件,其功能随所嵌入的工作流而变化。A large-scale, data-driven analysis of Jev projects collected from GitHub finds rapid early growth in Jev's public ecosystem, with both new projects and integration into existing repositories, and suggests that Jev serves as a reusable decision component whose functionality varies with the surrounding workflow.
Continnum,一种通过为 KV cache 保留引入 TTL 机制来优化多轮 Agent 工作负载任务完成时间的服务系统,与程序级 FCFS 结合时可保持多轮连续性,并降低 Agent 工作流的延迟。Continnum, a serving system to optimize job completion time for multi-turn agent workloads by introducing time-to-live mechanism for KV cache retention, and when combined with program-level first-come-first-serve preserves multi-turn continuity, and reduces delay for agentic workflows.
EngramRAG 是一种自适应记忆架构,将低延迟的 Waking State 反射与异步后台 Dreaming State 整合周期耦合,并引入融合密集向量、BM25 与 U-PPR 的三源混合检索,通过动态 Reciprocal Rank Fusion (RRF) 实现。The proposed EngramRAG is an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle, and introduces triple-source hybrid retrieval fusing dense vectors, BM25, and U-PPR via dynamic Reciprocal Rank Fusion (RRF).
TimeEvo 将 Agent 诊断出的失败聚类为能力缺口,为每个缺口规划测量,合成只产出证据的工具来填补缺口,并仅通过配对的准入门控接纳候选工具库。TimeEvo is proposed, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate.
D-JEPA 是一种决策对齐的潜在世界模型,从已执行结果中学习候选未来之间的决策相关关系,将决策相关关系结构确立为预测世界建模与有效控制之间的直接桥梁。D-JEPA is introduced, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes, establishing decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.
Just-In-Time Agent Memory (JAM) 是一个可训练框架,在运行时进行查询条件下的上下文构建,其任务性能优于 AOT 式记忆系统,同时比先前的可训练 Agent 记忆方法显著更高效。Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime, is proposed, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches.
Stashbird 是一种 Agent 记忆系统,通过显式来源溯源将源事件链接到派生记忆状态,在 LongMemEval-S 和 GroupMemBench 上准确率高于 Hindsight,在 EverMemBench 上与之相当。Stashbird is presented, an agent memory system that links source episodes to derived memory state through explicit provenance and achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.
本文提出 WideSWE,用于在跨仓库任务上评估编码 Agent,并在相同 prompt 下将其与联合执行进行对比,以考察逐仓库工作是否能缓解相关困难。This work introduces WideSWE to evaluate coding agents on cross-repository tasks, and compares it with joint execution under identical prompts to examine whether working on one repository at a time can alleviate difficulties.
提出 GAGAR,一个面向代码 agent RL 的质量感知信用再分配框架;结果表明,将基于测试的验证与成组的 agentic 评分相结合,可提升代码 agent RL 的质量与稳定性。GAGAR, a framework for quality-aware credit redistribution in code agent RL, is introduced and results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.
介绍 Relic,它将反复出现的协作失败转化为组织拥有的、可执行的协议,使协作经验成为超越其创建者仍持续有用的组织级持久状态。The introduction of Relic, which turns recurring collaboration failures into organization-owned, executable protocols, and how collaboration experience can become persistent organizational state that remains useful beyond the members who created it are shown.
提出反事实树上的注意力搜索(ASCT),将训练时的多步搜索转化为局部动作信用,将反事实评估与策略学习相连接,同时部署时只需使用 actor。Attentive Search over Counterfactual Trees (ASCT) is introduced, a framework that turns training-time multi-step search into local action credit and connects counterfactual evaluation to policy learning while deploying the actor alone.
提出 Context Language Models,能原生管理自身上下文,并可通过经标准 skill 优化循环演化的自然语言指令进行引导,在上下文管理任务上将未见数据的准确率最高提升 35.9 分,同时降低计算开销。Context Language Models are introduced, language models that natively manage their own context and can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute.
提出 CorpusMap,一个以语料中重复出现的实体为核心的导航层;这些实体可从文档自身识别,并能在不同来源间把单个文档与众多其他文档相连,表明实体可作为大型文档集合导航的有效锚点。CorpusMap is introduced, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources, suggesting that entities serve as effective anchors for navigating large document collections.
在所测试的每一对基座与指令微调模型中,指令微调都强化了模型对保留标记的偏好,且该差距在该通道上持续存在。In every base and instruction-tuned pair the authors test, instruction tuning strengthens the model's preference for reserved markers, and the gap persists on that channel.
KUPAS MASTER 是一个围绕九层认知语料构建的经验工程平台,将异构的工作记录与从业者访谈转化为 agent 可用的可追溯、可复用经验语料,提供了从个体隐性经验到组织知识与 agent 能力的可行路径。KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction, turns heterogeneous work records and practitioner interviews into traceable, reusable experience corpora for agents, and provides a practical path from individual tacit experience to organizational knowledge and agent capabilities.
本文提出一种异步 LLM 框架,允许用户(或 Agent 自身)定义具有重叠 memory 状态的推理协程,并展示 Qwen 3.x 模型无需任务专属训练即可在流式视频理解、电子游戏和监控中实现异步运行。This work develops an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states and showcases that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.