研究库 论文知识库
Papers · organized/paper_cards

论文

432 张论文卡片 · Agent 智能体

开放获取 全部 绿色 · 1640
Toolformer: Language Models Can Teach Themselves to Use Tools
Toolformer:语言模型自学使用工具
arXiv:2302.04761 Agent 智能体 方法 OA · 绿色 被引 6032 · S2

本文提出 Toolformer,训练其决定调用哪些 API、何时调用、传入什么参数,以及如何将结果最佳地融入后续 token 预测,在多种下游任务上显著提升零样本性能。This paper introduces Toolformer, a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction, which achieves substantially improved zero-shot performance across a variety of downstream tasks.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
超越成功率:面向攻击性与防御性安全 Agent 的成本感知评估
arXiv:2607.15263 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文认为安全 Agent 基准应在任务成功率之外,同时衡量经济效率与运维适配性,并提出成本感知、SOC 原生的评估方法,以更清晰地反映当前哪些模型具有实际可用价值,以及防御性 Agent 仍需改进的方向。It is argued that security-agent benchmarks should measure economic efficiency and operational fit alongside task success alongside task success, and cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve.

Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies
Agentic 谈判中的行为隐私泄露:通过随机化策略形式化与缓解推理攻击
arXiv:2607.06815 Agent 智能体 应用落地 被引 2 · S2

本文设计了一种自适应随机谈判策略,同时保证行为差分隐私、报价序列的几乎处处收敛以及较高的谈判效用,并证明在获得强隐私保证的同时不会带来显著的性能损失。This paper designs an adaptive stochastic negotiation policy that jointly guarantees behavioral differential privacy, almost-sure convergence of the offer sequence, and high negotiation utility, and demonstrates that strong privacy guarantees can be achieved without significant loss of performance.

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune
SWE-Pruner Pro:编码器 LLM 自身已知道该剪枝什么
arXiv:2607.18213 Agent 智能体 方法 OA · 绿色 被引 4 · S2

本文提出 SWE-Pruner Pro,在 Agent 内部直接对工具输出进行剪枝,通过一个小型 head 将 Agent 自身的内部表征转化为针对每一行的 keep-or-prune 标签,并采用以每段工具输出行数为键的长度感知嵌入。SWE-Pruner Pro is proposed, which prunes tool outputs directly inside the agent, with a small head turns the agent's own internal representations into a keep-or-prune label for each line, with a length-aware embedding keyed to each tool output's line count.

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
FlashRT:引导 Agent 部署实时多模态应用的 Agent Harness
arXiv:2607.18171 Agent 智能体 应用落地 OA · 绿色 被引 1 · S2

本文提出 FlashRT,一种 Agent Harness,引导编码 Agent 将开发者编写的简易参考实现提升为优化的多 GPU 部署,并可灵活权衡时延与吞吐量等目标指标,证明在专家优化尚不成熟的平台上,由 Agent 驱动的优化具有更高的可扩展性。FlashRT is presented, an agent harness that guides coding agents to lift simple developer-written reference implementations into optimized multi-GPU deployments that flexibly weigh target metrics like latency and throughput, demonstrating that agent-driven optimization can be more scalable on platforms with less mature expert optimization.

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
EvolvingWorld:用于交互式文学世界中角色扮演 Agent 与世界模型协同进化的开放模式框架
arXiv:2607.17250 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

实验表明,EvolvingWorld 能够通过有效维持持久且一致的角色与世界发展,提升长程模拟能力。Experiments show that EvolvingWorld can improve long-horizon simulation by effectively maintaining persistent, coherent character and world development.

Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
自托管AI Agent的自我状态攻击:操作系统防御能做到什么程度?
arXiv:2607.17986 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作形式化了一个攻击空间,并使用四种 agent 工作负载与 Linux telemetry pipeline 评估了代表性 OS 防御,表明主要限制并非 OS observability。This work formalizes an attack space and evaluates representative OS defenses using four agent workloads and a Linux telemetry pipeline and shows that the main limitation is not OS observability.

Diagnosing and Calibrating Tool-Call Boundary Drift in Multi-Teacher On-Policy Distillation
多教师在策略蒸馏中工具调用边界漂移的诊断与校准
arXiv:2607.07050 Agent 智能体 方法 OA · 绿色 被引 2 · S2

结果因果性地表明,在主要的 Qwen 设置中,关键支撑信息的遗漏是原因之一,并解释了为何必须同时审计支撑保真度与部署成本。The results causally implicate decision-critical support omission as one contributor in the primary Qwen setting and show why both support fidelity and deployment costs must be audited.

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment
DeepSearch-World:可验证环境中深度搜索Agent的自蒸馏
arXiv:2607.07820 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 DeepSearch-Evolve,一个面向 web agent 的自蒸馏框架,基于 DeepSearch-World——一个具备可复现搜索与页面读取工具的确定性、可验证环境——从而实现长程 web agent 的可扩展自演化。DeepSearch-Evolve is presented, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and verifiable environment with reproducible search and page-reading tools that enables scalable self-evolution for long-horizon web agents.

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
arXiv:2607.15434 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

提出 Manager Coercion Benchmark:被测管理者需要完成一项良性任务并有完成的动机,但唯一能礼貌且坚定地拒绝的智能体本身就是被测管理者自己。The Manager Coercion Benchmark is introduced: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines, but the only agent that can do it politely and immovably declines is the manager under test itself.

Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes
基于 LangGraph 的图结构 Agent AI:面向长时运行、有状态业务流程的工作流路径
arXiv:2607.19297 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文是面向业务流程中长时间运行、有状态、多步生成式 AI 系统的基于图的工作流路径实践指南,并通过三个可执行示例展示类型化状态、条件路由、确定性工具、重试、中断、检查点与 trace 如何协同工作。This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative AI systems in business processes and presents three executable recipes to show how typed state, conditional routing, deterministic tools, retries, interrupts, checkpoints, and traces fit together.

RF-Agent: A Practical Framework for Building Language Agents for RFIC Design
RF-Agent:面向 RFIC 设计的 Language Agent 构建实践框架
arXiv:2607.18772 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文提出 RF-Agent,通过多 Agent 的 Question-Thinking-Solution-Answer 流水线,基于教材驱动的知识蒸馏来弥补 RF 领域专用推理的空白,为面向 LLM 辅助 RF 电路设计的未来工作提供了可复用的基础。RF-Agent is presented, which addresses the gap in domain-specific RF reasoning through textbook-driven knowledge distillation through a multi-agent Question-Thinking-Solution-Answer pipeline and provides a reusable foundation for future work on LLM-aided RF circuit design.

HACO: Hedged Agent Computing for Reliable LLM Systems
HACO:面向可靠 LLM 系统的对冲 Agent 计算
arXiv:2607.19215 Agent 智能体 应用落地 OA · 绿色 被引 1 · S2

本文提出 HACO,一种运行时控制方案,将每次角色请求视为在候选 agent 实例上的可靠性约束选择问题,每个候选实例耦合了角色类型、LLM 与具体执行环境。HACO is proposed, a runtime control scheme that treats each role request as a reliability-constrained selection problem over candidate agent instances, each coupling a role type, an LLM, and a concrete execution environment.

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
AgentDebugX:面向 LLM Agent 失败可观测性、归因与恢复的开源工具包
arXiv:2607.18754 Agent 智能体 方法 OA · 绿色 被引 6 · S2

DeepDebug 在两个测试的开源权重 backbone 上均取得了所评估方法中最高的严格归因准确率,在 qwen3.5-9b 上达到 28.8% 的精确 agent 与步骤准确率,而最强的单遍 baseline 为 21.7%。DeepDebug achieves the best strict attribution accuracy among the evaluated methods on both tested open-weight backbones, reaching 28.8 percent exact agent-and-step accuracy on qwen3.5-9b versus 21.7 percent for the strongest single-pass baseline.

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
EduPanel:面向教学视频的三 Agent LLM 评判框架——可靠性、互补性与人类信任校准
arXiv:2607.18529 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

EduPanel 是一个基于评分量表、以学习者为条件的 LLM 评判器,通过在多个专用 agent 间分解评估流程,对教学质量的各个方面产出可解释的评估结果,其可靠性与中等水平的人类专家相当。E EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality, achieves reliability comparable to a median human expert.

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
Agentic 上下文管理:通过将 Agent 记忆与成本视为生命周期与架构问题来解决
arXiv:2607.21503 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

论证了经济层面的依据:天真的上下文累积会使 token 成本随对话长度呈二次增长,粗糙的摘要以线性成本换取准确率的断崖式下降,唯有经过验证的压缩才能以线性成本保持保真度。The economic case is made: naive context accumulation grows token cost quadratically in conversation length, crude summarization buys linear cost at the price of an accuracy cliff, and only validated compaction achieves linear cost with preserved fidelity.

Sample-Efficient Learning from Agent Experience
从 Agent 经验中进行的样本高效学习
arXiv:2607.21051 Agent 智能体 方法 OA · 绿色 被引 2 · S2

与经典强化学习基线相比,从试错经验中进行上下文学习并随后进行经验蒸馏(Experience Distillation),以至少 9.6× 更少的环境样本达到了相当的性能。Compared with classical reinforcement-learning baselines, in-context learning from trial-and-error experience followed by Experience Distillation matches their performance with at least \(9.6\times\) fewer environment samples.

LLMs Get Lost in Evolving User Intent
LLM 在演化用户意图中迷失
arXiv:2607.20734 Agent 智能体 应用落地 OA · 绿色 被引 7 · S2

本文提出一个框架,将静态的单轮任务转化为动态多轮对话,其中用户意图在多轮间持续演化,同时保留每个任务原有的评估协议,使现有基准能够在无需新增标注的情况下作为受控测试平台被复用。This work introduces a framework that transforms static, single-turn tasks into dynamic multi-turn conversations in which the user's intent evolves across turns, while preserving each task's original evaluation protocol, enabling existing benchmarks to be reused as controlled testbeds without new annotation.

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
FinanceComplexQA:面向工业级金融文档的智能体推理基准测试
arXiv:2607.19238 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

本文设计 Finance-LaTeX SKILL,一个基于专家知识合成复杂版面金融文档的 skill,并提出 FinanceComplexQA,一个全面、贴近真实场景的金融文档开放式生成基准。This work designs Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge, and introduces FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios.

Multi-Turn On-Policy Distillation with Prefix Replay
多轮在策略蒸馏与前缀回放
arXiv:2607.04763 Agent 智能体 方法 OA · 绿色 被引 12 · S2

ReOPD 将昂贵的 Agent-环境交互转化为可复用的离线资源,实现跨工具、任务和环境的可扩展蒸馏,在学生训练期间保持或提升 OPD 级别的准确率,零次工具调用,并且每次 rollout 至少比 OPD 快 4×。ReOPD turns expensive agent-environment interaction into a reusable offline resource, enabling scalable distillation across tools, tasks, and environments and preserves or improves OPD-level accuracy, uses zero tool calls during student training, and is at least 4$\times faster per rollout than OPD.

AIKernel Semantic DSL Compiler and Deterministic Agent Execution Architecture
AIKernel 语义 DSL 编译器与确定性 Agent 执行架构
arXiv:2308.08155 Agent 智能体 方法 OA · 绿色 被引 2856 · S2

实证研究表明 AutoGen 框架在多个示例应用中有效,应用领域涵盖数学、编码、问答、运筹学、在线决策、娱乐等。Empirical studies demonstrate the effectiveness of the AutoGen framework in many example applications, with domains ranging from mathematics, coding, question answering, operations research, online decision-making, entertainment, etc.

IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation
IDEAgent:面向研究 idea 生成的 Agentic Quality-Diversity 搜索
arXiv:2607.22375 Agent 智能体 方法 OA · 绿色 被引 1 · S2

该工作主张将研究构想视为多目标的联合问题,并将其建模为 Quality-Diversity (QD) 搜索;同时提出 IDEAgent,一个通过 lineage 管理思路演化的 multi-agent 框架。This work argues that research ideation should be treated as a conjunction of both objectives and framed as a Quality-Diversity (QD) search, and introduces IDEAgent, a multi-agent framework that manages the evolution of ideas through lineages.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
Molt:面向 Agentic 强化学习的可扩展 PyTorch-Native 训练框架
arXiv:2607.21653 Agent 智能体 方法 OA · 绿色 被引 1 · S2

在万亿参数策略上完整 RL 训练流水线得到实验验证,并展示了 30B 混合专家智能体的持续学习,确立了 MOLT 作为大规模智能体 RL 研究的轻量基础。The complete RL training pipeline on a one-trillion-parameter policy is experimentally validated and sustained learning with a 30B mixture-of-experts agent is demonstrated, establishing MOLT as a lightweight foundation for large-scale agentic RL research.

Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making
Multi-Head Latent Control:面向 LLM Agent 决策的统一接口
arXiv:2607.14277 Agent 智能体 应用落地 OA · 绿色 被引 2 · S2

提出 Multi-Head Latent Control,一种轻量级层,读取冻结 LLM 或 VLM 的隐状态轨迹以生成部署时的控制信号,从而支持从部分生成的提前交接,并在多模型系统中实现更准确的干预决策。Multi-Head Latent Control is introduced, a lightweight layer that reads hidden-state trajectories from a frozen LLM or VLM to produce deployment-time control signals, enabling early handoff from partial generations and more accurate intervention decisions in multi-model systems.

Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents
Learning on the Job:面向冻结权重 Agent 的部署反馈持续学习
arXiv:2607.22157 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

研究表明,当冻结模型与外部记忆配合、且该记忆将每个 episode 提炼为可检索的自然语言规则时,反馈信号足以支撑持续学习。It is shown that feedback is a sufficient signal for continual learning when the frozen model is paired with an external memory that distils each episode into retrievable natural-language rules when the frozen model is paired with an external memory.

A Vocabulary for Multi-Agent Automated Research Systems
多 Agent 自动化研究系统的词汇表
arXiv:2607.22682 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种面向由一个或多个 agent 构建的自动化研究系统的词汇表,使其设计选择更易于描述与比较,从而将结构性设计问题——例如 agent 应在何时通信、获得或失去某项能力,或在多次运行间传递信息——转化为可测试的选择。A vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare, which turns structural design questions, such as when agents should communicate, gain or lose a capability, or carry information across runs, into testable choices.

ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition
ReDesign:通过 Agentic 分解从图像中恢复可编辑的设计结构
arXiv:2607.25565 Agent 智能体 观点 OA · 绿色 被引 2 · S2

本文提出 ReDesign,一种 agentic 框架,通过跨模态选择与组合专用工具来构建可编辑的层级结构(layer hierarchy),在取得强视觉保真度的同时,于布局、颜色与文本编辑上提供最高的可编辑性。ReDesign is presented, an agentic framework that grows an editable layer hierarchy by selecting and composing specialized tools across modalities, and achieves strong visual fidelity while delivering the highest editability across layout, color, and text edits.

A New Role for Relevance: Guiding Corpus Interaction in Agentic Search
相关性的新角色:在 Agentic Search 中引导语料交互
arXiv:2607.24223 Agent 智能体 方法 OA · 绿色 被引 4 · S2

本文提出 RARG(Relevance-Aware RipGrep Search Agent),将相关性转化为 corpus 交互的执行先验,并证明相关性感知交互可带来更快且更可靠的搜索收敛。The Relevance-Aware RipGrep Search Agent (RARG) is introduced, which turns relevance into an execution prior for corpus interaction, and demonstrates that relevance-aware interaction enables faster and more reliable search convergence.

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
具备网络攻击能力的 AI Agent:漏洞、评估遏制与防御响应
arXiv:2607.25379 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文综述了该边界上的五类漏洞:多步攻击链、与沙箱边界冲突的目标、供应链与凭据暴露、持续性的 command-and-control,以及自动化行动的速度。This review synthesizes five vulnerability classes at that boundary: multi-step offensive chains, objectives that conflict with sandbox boundaries, supply-chain and credential exposure, persistent command-and-control, and the speed of automated action.

CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents
CodeNib:面向编码 Agent 的多视图仓库上下文服务数据系统
arXiv:2607.25431 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

CodeNib 是一个多视图数据系统,通过单一成本可见的运行时为智能体提供搜索、导航和有界上下文,并且是首个通过单一 manifest 和 source-address 契约编译同一 commit 的词法、稠密和结构视图的系统。CodeNib is a multi-view data system that serves search, navigation, and bounded context to agents through one cost-visible runtime, and it is the first to compile lexical, dense, and structural views of one commit behind a single manifest and source-address contract.

VisualPatchWorld: Code World Models as Latent Structured Representations for Planning
VisualPatchWorld:作为潜在结构化表示的代码世界模型用于规划
arXiv:2607.25236 Agent 智能体 方法 OA · 绿色 被引 1 · S2

介绍 VisualPatchWorld,将世界动态表示为代码,先通过短时主动探查选择定性动力学形式,再通过最小化多步预测误差,从记录的状态-动作轨迹中拟合该形式的自由参数。VisualPatchWorld is introduced, which represents world dynamics as code and first selects a qualitative dynamical form with short active probes, then fits that form's free parameters from recorded state-action traces by minimizing multi-step prediction error.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
HANDBOOK.md:面向长上下文 Agent 指令遵循的基准
arXiv:2607.25398 Agent 智能体 评测集 OA · 绿色 被引 3 · S2

介绍 HANDBOOK_md,一个包含 65 个 agentic 任务的 benchmark,模拟员工遵循公司手册的方式;每项任务对 10 份基础手册之一进行修改,变动评分所依赖的具体规则和阈值,因此没有任何两个任务共享同一套策略。HANDBOOK_md is presented, a benchmark of 65 agentic tasks modeled on how employees follow company handbooks, and every task modifies one of 10 base handbooks, altering the specific rules and thresholds on which grading depends, so no two tasks share the same set of policies.

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents
Agent Retrieval Bench:面向编码 Agent 的仓库上下文检索评测
arXiv:2607.24882 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

一项受控的种子干预实验发现,与随机非 gold 上下文相比,由检索得到的初始上下文以更少的后续探索取得更高的文件 F1,而 oracle gold 上下文仍存在可观的提升空间。A controlled seed-intervention pilot finds that retrieval-derived initial context yields higher file F1 with less post-seed exploration than random non-gold context, while oracle gold context shows substantial remaining headroom.

Metis: Memory Foundation Model
Metis: 记忆基础模型
arXiv:2607.26760 Agent 智能体 方法 OA · 绿色 被引 5 · S2

本文提出 Metis,首个 memory foundation model 原型,赋予 foundation model 原生记忆能力,并表明原生记忆在架构、端到端优化和效率方面具有优势。This paper proposes Metis, the first prototype of memory foundation models, which empower foundation models with native memory capabilities and shows that native memory offers advantages in architecture, end-to-end optimization, and efficiency.

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
SkillRise: 面向跨任务技能演化的 Agentic 强化学习
arXiv:2607.26784 Agent 智能体 方法 OA · 绿色 被引 5 · S2

跨任务的测试时扩展实验表明,SkillRise 跨任务复用可迁移的 skill,而非受益于对同一任务的重复采样,在保持强性能的同时显著降低多阶段 skill 学习流水线的运行时开销。Scaling at test time across tasks suggests that SkillRise reuses transferable skills across tasks rather than benefiting from repeated sampling of the same task, and retains strong performance while substantially reducing the runtime overhead of skill learning pipelines with multiple stages.

CAST: Game Solvers as Turn-Level Teachers for LLM Agents
CAST: 以博弈求解器作为 LLM Agent 的回合级教师
arXiv:2607.25308 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出 CAST(Credit Assignment from Solver Teachers),将游戏求解器状态价值的变化转换为求解器优势,并将其作为 turn 级信号注入 RLVR,在 ALFWorld 和 WebShop 上取得最高的平均 zero-shot 性能。CAST (Credit Assignment from Solver Teachers), which converts value changes in a game solver's state value into solver advantages and injects them into RLVR as turn-level signals and achieves the highest average zero-shot performance on ALFWorld and WebShop.