研究库 论文知识库
Papers · organized/paper_cards

论文

284 张论文卡片 · Agent 智能体 · 方法

开放获取 全部 绿色 · 1640
Structuring MoE Expert Selection for Agentic Reinforcement Learning
面向 Agentic 强化学习的 MoE 专家选择结构化
arXiv:2610.07332 Agent 智能体 方法

长程 LLM agent 常借助稀疏 mixture-of-experts(MoE)模型实现,但 agentic 行为与 MoE 结构的协同设计仍缺乏充分探索。本文系统研究 agentic 后训练与 MoE 专家选择之间的关联。在现成 MoE 模型中,我们观察到专家选择呈现出与 agentic 轨迹自然对齐的专门结构:agent 执行语义相似操作(如 READ、UPDATE)的回合之间,其专家路由的重叠度高于执行不同操作……Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing oper

ExperienceIndex: Artifact-Grounded Memory
ExperienceIndex:基于 artifact 的记忆
arXiv:2610.10091 Agent 智能体 方法

知识密集型任务需要通过推理共享的 artifact 语料库(如案件或科学文献)来回答多个问题。随着人类与这些语料库交互,他们自然地积累关于 artifact 的经验知识,从而能够快速识别每个新任务所需的完整相关 artifact 集合。然而,现有 AI agent 缺乏构建或复用此类 artifact-grounded 经验的合适记忆方案,导致答案质量降低且在线成本升高。现有记忆方案从先前任务求解过程中提取并复用信息……(原文截断)Knowledge-intensive tasks require answering many questions by reasoning about a shared corpus of artifacts (e.g., court cases, or scientific literature). As humans interact with these corpora, they naturally accumulate experiential knowledge about artifacts, enabling them to quickly identify the complete set of relevant artifacts for each new task. However, existing AI agents lack appropriate memory solutions to build or reuse such artifact-grounded experience, leading to lower answer quality and higher online cost. Existing memory solutions extract and reuse information from prior task-solvin

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery
从 Pareto 到偏好:通过摊销 agentic 策略发现实现个性化测试时扩展
arXiv:2610.09684 Agent 智能体 方法

测试时扩展(TTS)通过分配额外的推理算力来提升大语言模型的推理能力。现有提升 TTS 效率的方法通常一次针对单一资源维度优化准确率,推进 accuracy–cost 或 accuracy–latency 的 Pareto 前沿。然而用户需求是多维的:用户可能同时指定准确率、延迟与推理成本需求,而不同需求可能偏好不同的控制器。我们将个性化测试时扩展形式化为发现可执行控制器,以最大化……(原文截断)Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
SkillForge:通过动态技能生命周期协同演化技能与 agent
arXiv:2610.09832 Agent 智能体 方法

记忆增强强化学习增强了 LLM agent 解决复杂长程任务的能力。技能是此类记忆的一种形式,将指令与跨任务类型的适用条件配对。然而随着策略改进,若不加区分地保留所有技能,会导致过时或有害条目累积并误导 agent。我们提出 SkillForge,一种 agentic RL 方法,通过由适应性驱动的技能生命周期(试用、活跃、稳定、退役状态)来编译与演化技能库,使技能与模型在整个训练过程中协同演化。预 RL……(原文截断)Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A pre-RL

PhysEvo: Astra Can Act, Let It
PhysEvo:让 Astra 学会行动
arXiv:2610.08995 Agent 智能体 方法

Astra 具备行动能力,但可靠的操作取决于其观察和控制世界的系统。本文提出 PhysEvo,一个围绕单个冻结模型构建的物理递归自我改进(RSI)框架。任务 Agent 执行机器人任务,元 Agent 利用产生的轨迹诊断失败、修订工具与技能,并测试修正效果。元 Agent 还能改进自身的诊断工具,使保留的修订同时支持后续行动与后续自我改进。该过程发展出关节级控制、循证观察以及可复用的操作……Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipula

Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation
通过在线上下文蒸馏将 Agent 经验内化至扩散模型权重
arXiv:2610.07250 Agent 智能体 方法

将图像生成模型包裹在 Agent 框架中可有效提升 Text-to-Image 任务性能:该框架能利用记忆、技能、工作流编排、结果验证与迭代优化来持续构建和修订 prompt,从而生成更好的图像。然而这些增益对扩散模型而言是外部的,只有运行完整框架时才能实现。本文提出扩散在线上下文蒸馏(D-OPCD),将 Agent 改进后的 prompt 视为特权上下文,并将编码在 Agent 框架中的知识蒸馏进……Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby eliciting better images. These gains, however, remain external to the diffusion model and are realized only while the full harness runs. We propose Diffusion On-Policy Context Distillation (D-OPCD), which treats the agent-improved prompt as privileged context and distills the knowledge encoded in the agent harness into the

UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
UniSkill:为持续进化的策略学习与执行者对齐的技能提案
arXiv:2610.10164 Agent 智能体 方法

大语言模型 Agent 可通过保留从先前交互中提炼的可复用技能来跨任务提升能力。已有研究联合优化任务执行与技能提取,使策略与技能库协同进化。然而,随着执行者持续学习,通过技能在后续训练步骤中的复用对其进行奖励,可能将技能收益与执行者自身的改进混淆;而直接测试每个候选技能又需要代价高昂的额外执行者 rollout。本文提出 UniSkill,利用共享策略与环境交互并提出技能……Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions. Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve. However, as the actor continues learning, rewarding skill proposals through their reuse in subsequent training steps may conflate skill benefits with actor improvement, while directly testing each proposed skill requires costly additional actor rollouts. In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose ski

CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use
CADFather:通过协同工具调用实现自主 CAD 重建
arXiv:2610.09127 Agent 智能体 方法

从三维形状重建可编辑 CAD 模型仍是一项具有挑战性的工程任务。现有方法可以提出 CAD 操作,但没有任何单一提案源能对不同零件几何与不同重建阶段同样有效。本文介绍 CADFather,一个自主 Agent 系统,通过协调互补工具从三维网格恢复参数化 CAD 程序。视觉-语言助手检查目标与中间重建结果的渲染图像,再决定扩展哪些候选 CAD 程序、调用哪些工具、生成多少提案,以及……Reconstructing an editable CAD model from a 3D shape remains a challenging engineering task. Existing methods can propose CAD operations, but no single source of proposals works equally well across different part geometries and stages of reconstruction. We introduce CADFather, an autonomous agentic system that coordinates complementary tools to recover parametric CAD programs from 3D meshes. A vision-language assistant inspects renders of the target and intermediate reconstructions, then decides which candidate CAD programs to extend, which tools to invoke, how many proposals to generate, and

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
自回顾蒸馏:将事后经验转化为先验远见
arXiv:2610.08077 Agent 智能体 方法

可验证奖励的强化学习 (RLVR) 主要通过交互后的标量结果奖励将 Agent 经验转化为学习信号。然而对于组相对目标,当所有推演获得相同奖励时,该信号即消失,尽管这些轨迹可能包含关于任务内容及 Agent 失败方式的有用信息。我们提出一个互补问题:事后回顾能否教会 Agent 在行动前本可预见的内容?我们引入前瞻学习,利用事后经验从行动前的Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-

Agent Plasticity: Measuring Self-Improvement Through Experience
Agent 可塑性:通过经验衡量自我改进
arXiv:2610.08902 Agent 智能体 方法

AI Agent 日益在能够诊断失败并通过经验改进的环境中运行,然而现有评估主要衡量 Agent 在固定时间点能做什么,而非其学习效果如何。评估自我改进需要回答三个问题:未来性能是否改进并泛化至学习交互之外;新能力获取效率如何;自我改进过程在哪里失效?为回答这些问题,我们在受控环境中研究自我改进,其中 Agent 分摊AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future performance improve and generalize beyond the interactions that enabled learning; how efficiently are new capabilities acquired; and where does the self-improvement process break down? To answer these questions, we study self-improvement in a controlled setting where agents amortize pa

Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance
Inherit-MAS:通过工作流与执行继承实现多 Agent 系统的测试时演化
arXiv:2610.02396 Agent 智能体 方法

由大语言模型构建的多 Agent 系统 (MAS) 通过协调专业化 Agent 来处理复杂任务,但有效的工作流难以预先设计。测试时演化利用执行反馈来优化工作流,然而广泛的修订可能扰动有用组件,而重执行未变更的请求会带来冗余计算。受生物进化中继承与选择相互作用的启发,我们引入 Inherit-MAS,在工作流和执行层面显式化继承。元模型首先综合出由 worker Agent 组成的工作流Multi-agent systems (MAS) built from large language models coordinate specialized agents to tackle complex tasks, but effective workflows are difficult to design in advance. Test-time evolution refines workflows using execution feedback, yet broad revisions can disturb useful components, while re-executing unchanged requests can incur redundant computation. Inspired by the interplay of inheritance and selection in biological evolution, we introduce Inherit-MAS, which makes inheritance explicit at the workflow and execution levels. A meta-model first synthesizes a workflow of worker agents with

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
Memento 3:通过反思式规则手册实现基于模型的递归自我改进
arXiv:2610.11794 Agent 智能体 方法

在未知环境中学会行动需要智能体推断世界运作方式,并在新证据出现时修正理解。然而,有限的观察可能支持多个世界模型,它们都能解释过去的交互,但对未见状态有不同的预测。我们提出 Memento 3,在 Memento 系列的基础上,使冻结的 LLM 智能体能够通过外部记忆持续学习显式世界模型。智能体维护一个自然语言规则手册作为持久化语义记忆,记录可修正的环境动态假设,同时保留未知部分Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects

REMORY: Learning Residual Memory for Context Compaction
REMORY:学习用于上下文压缩的残差记忆
arXiv:2610.11287 Agent 智能体 方法

长时程智能体压缩其历史以在有限上下文窗口内继续执行,但单一的文本摘要可能无法支撑所有后续决策。我们提出 REMORY,一种神经记忆网络,通过有界序列的软记忆 token 来补充摘要。给定历史和摘要,该网络学习生成帮助冻结 LLM 近似其在完整历史下产生的续写内容的 token。这些 token 以摘要为条件并附加在其后,构成沿序列维度的残差连接类比。在 SummHay 上,REMORY 提升了Long-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the history and summary, the network learns to generate tokens that help a frozen LLM approximate the continuation it would produce with the full history. The tokens are conditioned on the summary and appended after it, forming an analogue of a residual connection along the sequence dimension. On SummHay, REMORY im

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
CanvasAgent:通过视觉工具编排实现复杂图像创建与编辑
arXiv:2607.05465 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出了用于复杂图像创建与编辑的大规模多模态工具调用数据集 CanvasCraft,以及通过多轮交互学习编排异构视觉工具的工具增强多模态 Agent CanvasAgent。CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and CanvasAgent, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction are introduced.

Multi-Turn Agentic Scientific Literature Search via Workflow Induction
基于工作流归纳的多轮智能体科学文献搜索
arXiv:2607.00597 Agent 智能体 方法 OA · 绿色 被引 2 · S2

结果表明,显式且可编辑的搜索工作流为将文献搜索智能体与复杂科学意图对齐提供了有效且可控的接口。The results show that explicit, editable search workflows provide an effective and controllable interface for aligning literature search agents with complex scientific intent.

A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory
A-TMA:解耦长时 Agent 记忆中的状态感知失效
arXiv:2607.01935 Agent 智能体 方法 OA · 绿色 被引 4 · S2

提出 ATMA:在现有记忆系统之上的状态感知叠加层,保留被替换记录与过渡记录,为查询所需的"目标状态视图"构建证据包,并向问答模块暴露当前、历史与过渡三类标签。This work proposes ATMA, a state aware overlay for existing memory systems, which keeps superseded and transition records in the bank, builds evidence packets for the query's requested state view, and exposes current, historical, and transition labels to QA.

AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
AgenticSTS:面向长时 LLM Agent 的有界记忆测试平台
arXiv:2607.02255 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出了一种 Agent 设计和一套经过验证、可复用的方法,用于研究显式记忆层如何影响长周期 LLM Agent 决策,并给出了一种替代的有界契约。An agent design and a validated, reusable methodology for studying how explicit memory layers shape long-horizon LLM-agent decisions, as well as an alternative bounded contract, are introduced.

TRACE: State-Aware Query Processing over Temporal Evidence Graphs for Conversational Data
TRACE:面向会话数据的时序证据图上的状态感知查询处理
arXiv:2607.00339 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 TRACE,一个针对演化会话数据在时序证据图上的查询处理框架,将词汇召回与证据重建分离,从而在长会话历史中实现有界的查询时推理。TRACE is presented, a query processing framework over temporal evidence graphs for evolving conversational data that separates lexical recall from evidence reconstruction, enabling bounded query-time reasoning over long conversational histories.

Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
将个性化建模为逆向规划:通过结构去噪学习潜在设计意图以实现 Agentic 幻灯片生成
arXiv:2607.00407 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文将 PSP 形式化为逆向规划问题,并提出 SPIRE,一个通过有意破坏干净幻灯片的视觉结构来近似求解 PSP 的原则性框架,从而构建一个可验证的去噪任务。This work formulates PSP as an inverse planning problem, and proposes SPIRE, a principled framework to solve PSP approximately, by intentionally corrupting the visual structures of clean slides, which creates a verifiable task to denoise the corruption.

Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming
保障 AI Agent 安全:面向多层 Agent 红队测试的统一框架
arXiv:2606.31227 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本文提出 AI-Infra-Guard,一个围绕单一观测组织 AI 红队测试的开源框架,是目前唯一覆盖所有层面(包括对日益扩展 AI Agent 能力的 Agent Skills 供应链审计)的开源框架。AI-Infra-Guard is presented, an open-source framework that organizes AI red teaming around a single observation, and is the only open-source framework to span all of these, including supply-chain auditing of the agent skills that increasingly extend AI agents.

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking
PixelEyes:解耦感知与推理以实现精准视觉证据定位
arXiv:2607.00115 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文探索多轮视觉推理,观察到 MLLM 反复无法定位目标,导致冗长的推理轨迹,进而提出 PixelEyes,一种将推理与感知显式解耦的多轮视觉推理 Agent。This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories, and proposes PixelEyes, a multi-turn visual reasoning agent that explicitly decouples reasoning from perception.

Agentic Abstention: Do Agents Know When to Stop Instead of Act?
Agentic Abstention:Agent 是否知道何时该停止而非行动?
arXiv:2606.28733 Agent 智能体 方法 OA · 绿色 被引 13 · S2

研究发现,模型规模、推理能力与 Agent 脚手架以不同方式影响弃答行为,能力更强或更大的模型有时反而在及时弃答上表现更差。It is found that model scale, reasoning, and agent scaffolding affect abstention in different ways, where larger or more capable models sometimes perform worse at timely abstention.

When Multi-Robot Systems Meet Agentic AI:Towards Embodied Collective Intelligence
当多机器人系统遇上 Agentic AI:迈向具身集体智能
arXiv:2606.27929 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文探讨了具身集体智能 (ECI) 这一未来多机器人范式,其中机器人团队将世界上下文、任务进度与技能经验作为共享资源进行累积与利用。This article explores Embodied Collective Intelligence (ECI), a future multi-robot paradigm in which a robot team accumulates and uses world context, task progress, and skill experience as shared resources.

From Detection to Action: Using LLM Agents for Fault-Tolerant Control
从检测到行动:基于 LLM Agent 的容错控制
arXiv:2606.28011 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文提出一种基于 Agentic Large Language Model (LLM) 的主动容错控制 (FTC) 框架,可将故障检测输出转化为基于特定工厂知识的、符合约束的恢复动作。该方法结合:(i) 将操作员职责分解为监测、规划、动作合成、仿真、验证与重新提示的多 Agent 工作流;(ii) 数字过程工厂孪生 (DPPT),提供工厂数据、模型以及用于执行前测试的仿真服务;(iii) 基于 CPSMod 本体构建的 Graph Retrieval-Augmented Generation (Graph RAG) 层。We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constraint-aware recovery actions grounded in plant-specific knowledge. The approach couples (i) a multi-agent workflow that decomposes operator duties into monitoring, planning, action synthesis, simulation, validation, and reprompting; (ii) a Digital Process Plant Twin (DPPT) that exposes plant data, models, and a simulation service for pre-execution testing; and (iii) a Graph Retrieval-Augmented Generation (Graph RAG) layer built on the CPSMod ontol

GBC: Gradient-Based Connections for Optimizing Multi-Agent Systems
GBC:用于多智能体系统优化的基于梯度的连接
arXiv:2606.28187 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Gradient-Based Connections(GBC)——一种面向多智能体系统的细粒度归因与优化方法,可提升多智能体性能,超越强力的单智能体与多智能体基线;且归因质量越高,优化效果越显著。Gradient-Based Connections (GBC) is proposed, an approach for fine-grained attribution and optimization of multi-agent systems that improves multi-agent performance and outperforms strong single-agent and multi-agent baselines and higher attribution quality is associated with greater optimization effectiveness.

OpenRCA 2.0: From Outcome Labels to Causal Process Supervision
OpenRCA 2.0:从结果标签到因果过程监督
arXiv:2606.27154 Agent 智能体 方法 OA · 绿色 被引 6 · S2

PAVE 是一种逐步式标注协议,利用来自故障注入的已知干预来重建因果传播路径;逐步式的因果真值正是可信的基于 LLM 的 RCA Agent 所缺失的关键一环。PAVE, a step-wise labeling protocol that leverages known interventions from fault injection to reconstruct causal propagation paths, is introduced, a step-wise causal ground truth is the missing piece for trustworthy LLM-based RCA agents.

To Run or Not to Run: Analyzing the Cost-Effectiveness of Code Execution in LLM-Based Program Repair
运行与否:分析基于 LLM 的程序修复中代码执行的成本效益
arXiv:2606.26978 Agent 智能体 方法 OA · 绿色 被引 1 · S2

一项关于 LLM 程序修复中执行行为的两阶段实证研究表明,当前 Agent 不加区分地使用执行,在收益甚微的实例上仍付出其代价,因此应被视为具有显式成本-收益权衡的资源。A two-stage empirical study of execution behavior in LLM-based program repair suggests that current agents apply execution indiscriminately, paying its cost on instances where it provides little benefit, and should be treated as a resource with an explicit cost-benefit tradeoff.

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
何时组合 LLM 更有帮助?——基于 67 个前沿模型的路由、投票与 Mixture-of-Agents 共失效上限研究
arXiv:2606.27288 Agent 智能体 方法 OA · 绿色 被引 11 · S2

路由、投票、级联、融合与 Mixture-of-Agents 等多模型 LLM 系统常被用于超越单模型精度;研究表明其增益受限于一个该领域鲜少报告的量化指标,且在缺乏强查询级路由信号时,组合模型很少能胜过单一最佳模型。Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy, it is shown that their gain is capped by a quantity the field rarely reports, and combining models rarely beats the single best model without a strong query-level routing signal.

Lifelong In-Context Learning with Transformers Requires Parametric Forms of Attention
基于 Transformer 的终身上下文学习需要注意力的参数化形式
arXiv:2606.25342 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文认为,将上下文学习(ICL)扩展至终身设置是 AI Agent 持续学习的实用方案;要在固定硬件预算下用 Transformer 理解终身上下文,需要注意力的参数化形式。It is argued that extending in-context learning to lifelong settings is a practical solution for continual learning in AI agents and that parametric forms of attention are needed to understand a lifetime of context with transformers on a fixed hardware budget.

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It
多步工具调用强化学习为何崩溃及监督信号如何修复
arXiv:2606.26027 Agent 智能体 方法 OA · 绿色 被引 4 · S2

研究发现,强化学习(RL)与监督微调(SFT)交错训练可显著提升稳定性,但在格式与内容分布外(OOD)评测下性能下降;并展示了多样化监督信号如何引导探索式学习。It is found that interleaving supervised fine-tuning with RL substantially improves stability, but exhibits degraded performance under format and content out-of-distribution (OOD) evaluation, and how diverse supervisory signals can guide exploratory learning is demonstrated.

SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History
SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History
arXiv:2606.08671 Agent 智能体 方法 OA · 绿色 被引 7 · S2

提出 SkillHone,一个基于持久决策历史实现 Agent Skill 持续进化的 harness;在内部工具辅助的分析场景中提升了准确率,并在未预先集成搜索栈的情况下优于商业支持的深度研究 Agent。SkillHone is introduced, a harness for continual agent skill evolution grounded in persistent decision history that improves accuracy on internal tool-mediated analysis scenarios and outperforms commercially backed deep-research agents without a pre-integrated search stack.

Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
arXiv:2511.07397 Agent 智能体 方法 OA · 绿色 被引 2 · S2

提出 conversational infill:让一个小型 talker 模型在外部 reasoner 模型产生结果前即时生成上下文相关的回复以掩盖延迟,并在推理过程中将 reasoner 流式输出的知识流畅地融合到回复中。Conversational infill is introduced, where a small talker model both immediately generates contextually grounded responses to hide the latency of an external reasoner model and fluently integrates streamed reasoner knowledge into its responses during inference.

PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents
PACMS:作为 LLM Agent 可插拔引擎的次模上下文选择
arXiv:2606.20047 Agent 智能体 方法 OA · 绿色 被引 1 · S2

对话式与工具使用的 LLM Agent 在上下文窗口中同时从多个方向被填充,而必须在多轮之间回忆信息的 Agent(即 memory 的典型场景)恰恰是 recency 截断失效的地方。Conversational and tool-using LLM agents operate over a context window that fills from several directions simultaneously, and agents that must recall information across many turns, the defining case for memory, are precisely where recency truncation fails.

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence
S-Agent:借助空间工具使用激发空间智能推理
arXiv:2606.20515 Agent 智能体 方法 OA · 绿色 被引 6 · S2

提出 S-Agent,一种面向连续多视图图像与视频理解与推理的空间工具使用 Agent 范式,以无需训练的方式持续提升开源与闭源 VLM。This work introduces S-Agent, a spatial tool-use agentic paradigm for understanding and reasoning over continuous multi-view images and videos, and consistently improves both open-source and closed-source VLMs in a training-free manner.

Probe-and-Refine Tuning of Repository Guidance for Coding Agents
仓库指导的探测-微调:用于编码 Agent
arXiv:2606.20512 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

揭示了指导的生成方式才是决定性变量,并提出 probe-and-refine tuning(探测-微调):通过合成 bug 修复探测任务,利用单次 LLM 调用迭代诊断并修补仓库的指导文件,调优过程中不涉及 Agent 循环或工具调用。It is shown that how the guidance is produced is the decisive variable, and probe-and-refine tuning is introduced, a procedure that uses synthetic bug-fix probes to iteratively diagnose and patch a repository's guidance file through single-shot LLM calls, with no agent loop or tool use during tuning.

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents
低权限即足够时:探究 LLM Agent 中过度特权的工具选择
arXiv:2606.20023 Agent 智能体 方法 OA · 绿色 被引 11 · S2

提出一种特权感知的训练后防御方法,教导 Agent 优先选用足够的低权限工具,仅在必要时升级;该方法在保留通用能力的同时大幅减少了不必要的高权限工具使用。A privilege-aware post-training defense that teaches agents to prefer sufficient lower-privilege tools and escalate only when necessary is introduced, showing that this defense substantially reduces unnecessary high-privilege tool use while preserving general capabilities.