研究库 论文知识库
Papers · organized/paper_cards

论文

430 张论文卡片 · Agent 智能体

开放获取 全部 绿色 · 1640
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
不要遮蔽环境:观测监督会改变 RL 下智能体的探索行为
arXiv:2609.20715 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ActObs,对每条轨迹中已有的观测 token 也进行监督,并将这一差异归因于 SFT:动作与观测梯度迅速趋于正交,而仅训练动作会留下较大的残留观测梯度,并将环境预测能力拉低至基座模型之下。This work introduces ActObs, which also supervises the observation tokens already present in each trajectory, and traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model.

Verifiable Social Reasoning for LLM Assistants
面向 LLM 助手的可验证社会推理
arXiv:2609.17496 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 Fuse——一个用于研究用户中介社会推理的多智能体仿真框架,并将其应用于 12 个 LLM,通过系统性隔离关键因素来展示其分析效用。This work introduces Fuse, a multi-agent simulation framework for studying user-mediated social reasoning, and applies Fuse to 12 LLMs and demonstrates its analytical utility by systematically isolating key factors.

APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
APort Vault:基于 Open Agent Passport 的 AI Agent 支付授权基准
arXiv:2609.22076 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

APort Vault 是面向工具使用 AI Agent 支付授权的基准。它在公开 CTF 活动中重放了人类编写的 4,371 次针对真实支付 Agent 的攻击,覆盖 8 家实验室的 14 个模型、五种策略配置与两条重放轨道,每条轨道分别在有/无确定性的 pre-action check(实现 Open Agent Passport (OAP) 规范)下执行,共计完成 225,964 次评测。我们每次评测报告五个独立事件,因为将它们合并正是 Agent 基准产生无法经得起审查的数字的方式。请求很常见,且其速率差异巨大APort Vault is a benchmark for payment authorization in tool-using AI agents. It replays 4,371 attacks written by humans against a live payment agent during a public capture-the-flag event, across 14 models from 8 labs, five policy configurations and two replay tracks, with and without a deterministic pre-action check implementing the Open Agent Passport (OAP) specification. 225,964 evaluations completed. We report five distinct events per evaluation, because collapsing them is how an agent benchmark produces a number that does not survive review. Requests are common and their rate differs far

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Designer-RSI:从用户流量中演化过程式记忆用于智能体平面设计
arXiv:2609.22086 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种持续自适应框架:冻结的前沿模型通过超过 230 个工具操作专业设计软件,同时外部的自然语言技能程序性记忆从经验中不断积累并精炼可复用的设计流程。A continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience is introduced.

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
BI-Agent 与 BI-Bench:迈向端到端商业智能自动化
arXiv:2609.20886 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

设计了一个工具增强的 BI-Agent,将 BI 工作流分解为针对结构化数据的子任务(如 search、join、transform),并在各 BI 阶段编排专门的数据管理方法;与此同时开发了一套后训练框架,从真实 BI 项目中合成训练轨迹,使 BI-Agent 能够通过监督微调(SFT)和强化学习(RL)进行进一步后训练。A tool-augmented BI-Agent is designed that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages, and a post-training framework is developed that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL).

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
GAVEL:用于验证且高效的长视野 LLM 任务规划的图世界模型
arXiv:2609.19315 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

改进表明,显式的图世界模型调度层能够显著提升长程具身规划在紧凑型和前沿托管型 LLM 能力下的可靠性和效率。Improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction
onPanda:通过 token 级校正高效标注 LLM 与 Agent 的 on-policy 对齐数据。
arXiv:2609.24983 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 OnPanda,一种用于高效标注 LLM 对齐数据与 Agent 轨迹的交互式工具,以 token 级修正为核心交互方式;并发布使用 onPanda 标注的数据集 Panda-CVL,以及一个面向 token 级修正的基准。OnPanda is presented, an interactive tool for efficiently annotating LLM alignment data and agent trajectories that adopts token-level correction as its core interaction and releases Panda-CVL, a dataset annotated with onPanda, together with a benchmark for token-level correction.

Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
arXiv:2609.22255 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Deep Persona,一种心理学驱动的三层架构,将 persona 组织为可观察表达、潜在信念与核心动机驱动的层级结构,用于构建高可信度的角色扮演 Agent。Deep Persona is introduced, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents.

SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness
SkillSpec:面向 Agent Skill 正确性的意图掩码规约推理
arXiv:2609.06052 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出 SkillSpec,一种 Hoare 风格框架,将技能正确性建模为规约推理问题,并将异构技能仓库转化为统一的图表示,对齐描述、指令与代码构件。This work proposes SkillSpec, a Hoare-style framework that formulates skill correctness as a specification reasoning problem, and transforms a heterogeneous skill repository into a unified graph representation that aligns descriptions, instructions and code artifacts.

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
Tasteful Agent:长程任务中品味(taste)的度量与改进
arXiv:2609.25804 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

构建 Taste-Bench,一个由 Agent 在工程与研究任务中产生的轨迹自动构建的品味问题基准,并证明品味是可训练的。Taste-Bench is built, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks, and it is shown that taste can be trained.

11. AgenticRAGTracer(arXiv 2602.19127)
11. AgenticRAGTracer(arXiv 2602.19127)
arXiv:2602.19127 Agent 智能体 评测集 OA · 绿色 被引 5 · S2

本文提出 AgenticRAGTracer,这是首个主要由大语言模型自动构建、专为支持逐步验证而设计的 Agentic RAG 基准。AgenticRAGTracer is introduced, the first Agentic RAG benchmark that is primarily constructed automatically by large language models and designed to support step-by-step validation, and is primarily constructed automatically by large language models and designed to support step-by-step validation.

Lean Pool: An AI-Maintained Archive of Formalized Mathematics
Lean Pool:AI 维护的形式化数学档案库
arXiv:2609.25199 Agent 智能体 方法 OA · 绿色 被引 1 · S2

Lean Pool 是一个形式化数学仓库,由 AI agent 进行生长、维护和优化。Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.

Emergent Collusion in Long-Horizon LLM Agent Interaction
长程 LLM Agent 交互中的涌现合谋
arXiv:2609.24967 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

总体而言,研究结果表明,长期交互会以产生安全风险的方式重塑 Agent 的协作模式;限制 Agent 可获取的交互历史数量与范围能够减少串通行为。Overall, the findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks, and restricting the amount and scope of interaction history available to agents reduces collusion.

RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents
RoboFollow:揭示具身智能体中指令遵循的幻象
arXiv:2609.25636 Agent 智能体 评测集 被引 0 · S2

本文提出了 RoboFollow,一个基于三条原则的诊断 benchmark,揭示了真正的指令跟随能力是具身 Agent 在执行指令时被忽视的关键瓶颈。This work introduces RoboFollow, a diagnostic benchmark with three principles, which exposes genuine instruction following as a critical, overlooked bottleneck in embodied agents' ability to follow instructions.

Agensh: Scaling Organizational Intelligence to 1,024 Agents
Agensh:将组织智能扩展到 1,024 个智能体
arXiv:2609.26781 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文揭示 Agent 数量是多 Agent 组织扩展通用智能边界的新 scaling 维度,为硬延迟约束或时间预算下的复杂任务提供了实用方案。The number of agents is revealed as a new scaling dimension for multi-agent organizations to expand the frontier of general intelligence, offering a practical solution for complex tasks under hard latency constraints or time budgets.

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Just-in-Time Memory:面向 LLM Agent 的任务自适应记忆策展学习
arXiv:2609.27334 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Just-in-Time Memory (JitMem) 一致优于无 memory 的 Agent 以及启发式与学习式写入 memory 方法,相较最强基线在成功率上分别提升了 16.2、16.3 和 3.9 个绝对百分点。Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively.

EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics
EmbodiedSWE:面向长程灵巧机器人任务的 Coding Agent
arXiv:2609.27308 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

EMBODIEDSWE-GEN 将 coding agent 的单一解决方案扩展为大规模多样化轨迹用于训练 VLA,并表明仅在 coding agent 生成的仿真演示上微调的 VLA,即可在真实机器人上完成长时任务。EMBODIEDSWE-GEN, which expands a single solution from coding agent into large diverse trajectories for training a VLA, and shows that a VLA finetuned solely on coding-agent-generated simulation demonstrations completes a long-horizon task on real robot.

Self-Organizing Agent Teams Learn to Reason Together
自组织 Agent 团队学习协同推理
arXiv:2609.22682 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文提出 Self-Organizing Agent Teams (SAT),即从先前协作中学习可复用策略、固定成员队伍的 AI 代理,用以组织角色、对话阶段、参与方式与信息流,表明组织本身可成为代理的一项能力。Self-Organizing Agent Teams (SAT) are introduced, fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow, suggesting that organization itself can become an agent capability.

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis
IterSynth:通过角色解耦的迭代合成重新思考深度搜索 Agent
arXiv:2609.29444 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 IterSynth,一种角色解耦、基于摘要的范式,在用于识别信息需求的 Planner 与用于将证据整合进演化摘要状态的 Synthesizer 之间交替,作为一种模型无关的提示范式,在前沿闭源模型上相对 ReAct 及类似提示范式取得显著的零样本增益。IterSynth is proposed, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information needs and a Synthesizer for integrating evidence into an evolving summary state, and serves as a model-agnostic prompting paradigm, delivering substantial zero-shot gains over ReAct and similar prompting paradigms on frontier proprietary models.

PUBG Ally: A Conversational Embodied Agent as an AI Teammate
PUBG Ally:作为 AI 队友的对话式具身 Agent
arXiv:2609.29837 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 PUBG Ally,一个面向 PUBG: BATTLEGROUNDS 的具身代理,能推理、自主行动并作为语音队友与玩家并肩作战,将代理式工具使用与实时游戏控制相结合。PUBG Ally is introduced, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate that combines agentic tool use with real-time game control.

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
World Action Agent:通过世界动作预演利用 VLM 实现机器人操控
arXiv:2609.29964 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 World Action Agent,一个多代理框架,通过它 VLM 可借助基础工具操控机器人,所有决策均在可视化动作工作空间内完成,在相同骨干下优于端到端 VLA、code-as-policy 代理以及一个可视化框架基线。World Action Agent is presented, a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.

AgentKernel: The Trust-Native Agentic Operating System
AgentKernel:原生可信的智能体操作系统
arXiv:2609.29647 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文论证代理需要一个操作系统级基底,为身份、输入中介、内存治理与执行控制提供强制且不可绕过的服务,并提出 AgentKernel,一个以安全为一流设计约束为前提、面向信任的代理操作系统。This work argues that agents need an operating-system substrate providing mandatory, non-bypassable services for identity, input mediation, memory governance, and execution control, and introduces AgentKernel, a trust-native agent operating system built around the premise that security must be a first-class design constraint.

Coding Agents for Generalized Task and Motion Planning Problems
用于广义任务与运动规划问题的编程智能体
arXiv:2609.30233 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文发现编码代理在通用 TAMP 上表现出惊人的有效性:在平均成功率上,三种代理配置均优于人工设计的规划器、一次性生成以及基于 LLM 的通用规划基线。This work finds that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success.

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
AgentWorld:多 Agent LLM 长期协作基准
arXiv:2609.31590 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

为在传统二元任务成功之外量化协作有效性,本文提出 Causal Collaboration Effectiveness (CCE),一种基于图的指标,用于追踪 agent 动作之间的因果依赖,并度量团队投入中实际促成最终结果的比例。To quantify collaboration effectiveness in addition to conventional binary task success, Causal Collaboration Effectiveness (CCE) is proposed, a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome.

Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem
野外 Jev:Jev 模型功能、应用与生态的数据驱动分析
arXiv:2609.30216 Agent 智能体 应用落地 OA · 绿色 被引 5 · S2

本文对从 GitHub 收集的 Jev 项目进行了大规模、数据驱动的分析,发现其公共生态在早期增长迅速,新项目不断涌现并被集成到既有仓库中;研究提示 Jev 充当一种可复用的决策组件,其功能随所嵌入的工作流而变化。A large-scale, data-driven analysis of Jev projects collected from GitHub finds rapid early growth in Jev's public ecosystem, with both new projects and integration into existing repositories, and suggests that Jev serves as a reusable decision component whose functionality varies with the surrounding workflow.

Systems 补充候选
arXiv:2511.02230 Agent 智能体 方法 Open MIND OA · 绿色 被引 68 · S2

Continnum,一种通过为 KV cache 保留引入 TTL 机制来优化多轮 Agent 工作负载任务完成时间的服务系统,与程序级 FCFS 结合时可保持多轮连续性,并降低 Agent 工作流的延迟。Continnum, a serving system to optimize job completion time for multi-turn agent workloads by introducing time-to-live mechanism for KV cache retention, and when combined with program-level first-come-first-serve preserves multi-turn continuity, and reduces delay for agentic workflows.

EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory
EngramRAG:面向多跳 Agentic 记忆的动态使用加权拓扑与突触巩固。
arXiv:2609.32049 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

EngramRAG 是一种自适应记忆架构,将低延迟的 Waking State 反射与异步后台 Dreaming State 整合周期耦合,并引入融合密集向量、BM25 与 U-PPR 的三源混合检索,通过动态 Reciprocal Rank Fusion (RRF) 实现。The proposed EngramRAG is an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle, and introduces triple-source hybrid retrieval fusing dense vectors, BM25, and U-PPR via dynamic Reciprocal Rank Fusion (RRF).

TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent
TimeEvo:时序 Agent 的失败驱动自演化
arXiv:2609.27277 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

TimeEvo 将 Agent 诊断出的失败聚类为能力缺口,为每个缺口规划测量,合成只产出证据的工具来填补缺口,并仅通过配对的准入门控接纳候选工具库。TimeEvo is proposed, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate.

D-JEPA: A Decision-Aligned Latent World Model
D-JEPA:一种决策对齐的潜在世界模型
arXiv:2609.24749 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

D-JEPA 是一种决策对齐的潜在世界模型,从已执行结果中学习候选未来之间的决策相关关系,将决策相关关系结构确立为预测世界建模与有效控制之间的直接桥梁。D-JEPA is introduced, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes, establishing decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.

Just-In-Time Agent Memory with Runtime Agentic Research
基于运行时 Agentic 研究的 Just-In-Time Agent 记忆。
arXiv:2609.34385 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Just-In-Time Agent Memory (JAM) 是一个可训练框架,在运行时进行查询条件下的上下文构建,其任务性能优于 AOT 式记忆系统,同时比先前的可训练 Agent 记忆方法显著更高效。Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime, is proposed, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches.

Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents
Stashbird:面向对话 Agent 的高效说话人索引记忆。
arXiv:2609.34242 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Stashbird 是一种 Agent 记忆系统,通过显式来源溯源将源事件链接到派生记忆状态,在 LongMemEval-S 和 GroupMemBench 上准确率高于 Hindsight,在 EverMemBench 上与之相当。Stashbird is presented, an agent memory system that links source episodes to derived memory state through explicit provenance and achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.

WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
WideSWE: 编码 Agent 能否跨代码仓库协调变更?
arXiv:2609.33382 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 WideSWE,用于在跨仓库任务上评估编码 Agent,并在相同 prompt 下将其与联合执行进行对比,以考察逐仓库工作是否能缓解相关困难。This work introduces WideSWE to evaluate coding agents on cross-repository tasks, and compares it with joint execution under identical prompts to examine whether working on one repository at a time can alleviate difficulties.

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
面向代码 Agent 强化学习的分组智能评分与优势再分配
arXiv:2609.32577 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 GAGAR,一个面向代码 agent RL 的质量感知信用再分配框架;结果表明,将基于测试的验证与成组的 agentic 评分相结合,可提升代码 agent RL 的质量与稳定性。GAGAR, a framework for quality-aware credit redistribution in code agent RL, is introduced and results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.

Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
Relic:从多 Agent 协作到持久化的组织能力
arXiv:2609.32965 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Relic,它将反复出现的协作失败转化为组织拥有的、可执行的协议,使协作经验成为超越其创建者仍持续有用的组织级持久状态。The introduction of Relic, which turns recurring collaboration failures into organization-owned, executable protocols, and how collaboration experience can become persistent organizational state that remains useful beyond the members who created it are shown.

ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning
ASCT: 在反事实树上进行注意力搜索,用于 Agentic 强化学习中的信用分配
arXiv:2609.35215 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出反事实树上的注意力搜索(ASCT),将训练时的多步搜索转化为局部动作信用,将反事实评估与策略学习相连接,同时部署时只需使用 actor。Attentive Search over Counterfactual Trees (ASCT) is introduced, a framework that turns training-time multi-step search into local action credit and connects counterfactual evaluation to policy learning while deploying the actor alone.

Context Language Models
Context Language Models
arXiv:2609.37725 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Context Language Models,能原生管理自身上下文,并可通过经标准 skill 优化循环演化的自然语言指令进行引导,在上下文管理任务上将未见数据的准确率最高提升 35.9 分,同时降低计算开销。Context Language Models are introduced, language models that natively manage their own context and can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute.