研究库 论文知识库
Papers · organized/paper_cards

论文

284 张论文卡片 · Agent 智能体 · 方法

开放获取 全部 绿色 · 1640
DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents
DianShi-RxnDB:面向研究者和 AI Agent 的、完全自动化构建的大规模细粒度有机反应数据平台
arXiv:2609.06703 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 DianShi-RxnDB,一个通过全自动抽取与归一化流水线(整合专利文本、图像和反应路线图)构建的大规模细粒度有机反应数据平台。D DianShi-RxnDB is presented, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes integrating patent text, images, and reaction schemes.

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
AgentGrad:面向多 Agent 系统的干预引导 prompt 优化
arXiv:2609.08572 Agent 智能体 方法 OA · 绿色 被引 2 · S2

提出 AgentGrad,一种基于序贯干预与语义文本梯度抽象的多智能体系统 prompt 优化框架,在 5 个 MAS benchmark 上取得 SOTA 性能,同时降低优化耗时与成本AgentGrad is proposed, a prompt optimization framework for multi-agent systems based on sequential intervention and semantic textual gradient abstraction that achieves state-of-the-art performance across five MAS benchmarks while reducing wall-clock optimization time and optimization cost.

PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
PARSER:长上下文 LLM Agent 的并行读取与深度推理
arXiv:2609.06702 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文介绍 PARSER,将阅读与推理解耦,对证据位置、顺序与距离的扰动具有鲁棒性——这些条件会导致序列方法产生大幅精度波动——同时将推理延迟降低多达 11 倍。PARSER, which decouples reading from reasoning, is introduced, which is robust to perturbations in evidence position, order, and distance, conditions that cause large accuracy swings in sequential methods, while reducing inference latency by up to 11x.

Memory Compression for High-Fanout Agent Sandboxes
面向高扇出 Agent Sandboxes 的 Memory 压缩。
arXiv:2609.11294 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文介绍 AgentZip,第一个专为 AI Agent 沙箱设计的内存压缩系统,将压缩范围扩展到任何具有收益表示的页面,并将开销控制从压缩时页面选择转移到恢复时预取。AgentZip is presented, the first memory compression system designed specifically for AI-agent sandboxes, which broadens the compression scope to any page with a profitable representation and shifts overhead control from compression-time page selection to restore-time prefetching.

But How Would AI Agents Run a Town's Economy?
AI Agent 究竟要如何运行一座城镇的经济?
arXiv:2609.11108 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在真实的博卡拉湖畔地理环境中,由 100 个配备记忆机制的大语言模型 Agent 管理一个封闭且守恒的空间经济,并运行该多 Agent 模拟长达 26 个模拟周,远超典型 Agent 社会研究 1–2 周的时长。100 memory-equipped large language model agents in charge of a closed, money-conserving spatial economy on real Pokhara Lakeside geography and ran this multi-agent simulation for up to 26 simulated weeks, well past the 1-2 weeks typical of agent-society studies.

Memory as Plans: World-Action Modeling with Memory-Grounded Planning
以记忆为规划:基于记忆锚定的世界-动作建模规划
arXiv:2609.11561 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MaP-WAM,一个将记忆作为规划的 Memory-as-Plans 框架,将依赖记忆的世界动作建模分解为基于记忆的规划和以规划为条件的执行,并以长期多模态情景上下文作为规划时证据,而非反复对执行器输入完整历史。MaP-WAM is introduced, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history.

DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat
DRG-MAPPO:面向协同空战的多智能体强化学习层次动态角色图
arXiv:2609.11155 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

实验结果表明 DRG-MAPPO 达到了 87% 的 SOTA 胜率,表明该框架在合作空战中有效平衡了关系建模、可解释性和优化稳定性。Experimental results demonstrate that DRG-MAPPO achieves a state-of-the-art win rate of 87%, suggesting that the framework effectively balances relational modeling, interpretability, and optimization stability for cooperative air combat.

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
COBRA-Skills:基于上下文 Bandit 引导进化的 Agent 技能优化
arXiv:2609.11682 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

COBRA-Skills 是一个高效框架,将技能优化建模为在动态演化的候选空间上的预算式序贯优化,对 Agent harness 变化保持鲁棒,并在目标模型自身用于技能生成与优化时依然有效。COBRA-Skills is introduced, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space and remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.

Online Learning with LLM Experts from Limited Feedback
基于有限反馈的 LLM 专家在线学习
arXiv:2609.05820 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本研究探索在反馈有限的在线场景下,将 prompt 自适应路由到大语言模型专家以最大化响应质量,并提出了策略性地选择和观察奖励以最小化遗憾的算法。This work studies adaptive routing of prompts to large language model experts to maximize response quality in an online setting with limited feedback and proposes algorithms that strategically select and observe rewards to minimize regret.

Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
超越 Top-k 技能检索:面向 LLM Agent 的多样性感知技能路由
arXiv:2609.05824 Agent 智能体 方法 OA · 绿色 被引 2 · S2

提出了 Diverse Skill Routing,一个具备多样性感知能力的重排序框架,使用 Determinantal Point Process 在相关性与非冗余性之间取得平衡,在强 pointwise 重排序基线之上提升了召回率与完整覆盖率,且在多技能 query 上增益更大。Diverse Skill Routing is proposed, a diversity-aware reranking framework that uses a Determinantal Point Process to balance relevance and non-redundancy and improves recall and full coverage over a strong pointwise reranking baseline, with larger gains on multi-skill queries.

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work
Occamy-1.0:面向协作的开放帕累托前沿 35B 智能体
arXiv:2609.11977 Agent 智能体 方法 OA · 绿色 被引 3 · S2

提出成本高效的协同模型 Occamy-1.0,由后训练 checkpoint Qwen3.6-35B-A3B 继续训练得到,位于所观测成本-性能帕累托前沿的低成本拐点处。Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint, is presented and placed at the low-cost knee of the observed cost--performance Pareto frontier.

RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
RSIAgent:在新环境中通过递归自我改进实现自主探索
arXiv:2609.15364 Agent 智能体 方法 OA · 绿色 被引 3 · S2

提出 RSIAgent,一种无需训练、通过自主构建记忆实现递归自我改进的多 Agent 框架,显著增强了强开源模型,使 Kimi-K3 与 GLM-5.3 超越包括 GPT-6.3 在内的前沿闭源模型。RSIAgent is introduced, a training-free multi-agent framework for recursive self-improvement through autonomous memory construction that substantially improves strong open-source models, enabling Kimi-K3 and GLM-5.3 to outperform frontier closed-source models including GPT-6.3.

HazardAuditor: From Executable Threats to Safer Computer-Use Agents
HazardAuditor:从可执行威胁到更安全的 Computer-Use Agent
arXiv:2609.15134 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 HazardAuditor,一个执行驱动的框架,在受控环境中运行异构 Agent 并将其交互归一化为规范事件表示以实现跨框架监督;观察到 token 级后训练目标与生成式守卫存在结构性失配,导致更长的推理链主导梯度更新。HazardAuditor is introduced, an execution-grounded framework that runs heterogeneous agents in controlled environments and normalizes their interactions into a canonical event representation for cross-framework supervision, and observes that token-level post-training objectives create a structural mismatch for generative guards, causing longer rationales to dominate gradient updates.

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
当 Agent 变慢:通过 Elo-per-token 分析理解 LLM Agent 的测试时策略
arXiv:2609.15309 Agent 智能体 方法 OA · 绿色 被引 2 · S2

将规模化拐点定义为边际 Elo 增益等于独立采样参考时的单次会话预算;提出 Elo-per-token 分析,跟踪每个 token 预算下找到的最优解,并使用 Bradley-Terry 模型将各任务内的排序聚合为跨不同评分尺度任务的 Elo 评分。This work defines the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference, and proposes Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales.

Enabling Creative Exploration for Vibe Design Agents
为 Vibe Design Agent 启用创造性探索
arXiv:2609.15078 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出结构化设计规范作为探索替代 UI 概念的实用控制点,同时通过将设计方向显式化为中间决策的推理架构,保持下游生成设置固定。This work identifies structured design specifications as a practical control point for exploring alternative UI concepts while keeping downstream generation settings fixed through an inference architecture that makes design direction an explicit intermediate decision.

Agent as Policy for Robotic Manipulation
Agent作为机器人操作策略
arXiv:2609.12541 Agent 智能体 方法 OA · 绿色 被引 7 · S2

本工作证明通用 agent 可在整个任务执行过程中直接驱动物理机器人,无需任何任务或环境专属训练;并提出 Agent as Policy(AGP),将任务规划与执行置于 agent 的控制之下This work demonstrates that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training and introduces Agent as Policy (AGP), which places task planning and execution under the agent's control.

The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
人类构建的最后一个AI:迈向真正的递归自我改进
arXiv:2609.11873 Agent 智能体 方法 OA · 绿色 被引 4 · S2

该工作利用 Headroom-Closed Index(HCI)揭示现有 LLM 的问题,并提出 RSI 概念及其发展路线图:从改进执行自主性、改进策略自主性、经验获取自主性、环境适应自主性,到递归元改进。This work uses the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, and introduces the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement.

Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
广义 Agent 迭代:迭代式策略改进与递归自改进的统一形式化框架
arXiv:2609.13406 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Generalized Agent Iteration (GAI),一个形式化框架,将迭代策略改进和 RSI 描述为同一学习范式的两种情形,基于经典理论,使现有系统可比较,并为分析和设计新系统提供原则性基础。This paper proposes Generalized Agent Iteration (GAI), a formal framework that describes iterative policy improvement and RSI as two cases of a single learning paradigm that rests on the classical account, makes existing systems comparable, and provides a principled basis for analyzing and designing new ones.

Agora: Git as Shared Memory for Collective AutoResearch
Agora: 以 Git 作为集体 AutoResearch 的共享记忆
arXiv:2609.18094 Agent 智能体 方法 OA · 绿色 被引 2 · S2

报告了一次持续近 12 天的运行:13 个 LLM worker 在无任务分配、无中央规划器的情况下,使用 Agora 解决了一个权重迁移问题,并记录了 agent 如何复用与验证共享工作。A run of nearly 12 days is reported in which 13 language-model workers, with no assigned tasks or central planner, used Agora to solve a weight-transfer problem and documents how agents reused and verified shared work.

HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses
HypoEvolve:遗传算法赋能多智能体 LLM 发现科学假设
arXiv:2609.15938 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出一种代际遗传算法,用以协调专门的 LLM Agent,整合机制性论证、重新审视假设并评估证据与可检验性,推动了自主科学的愿景——AI 研究团队实现超越单个模型的发现能力。This work proposes a generational genetic algorithm to coordinate specialized large language model agents that integrate mechanistic arguments, reconsider assumptions, and assess evidence and testability that advances a vision of autonomous science in which AI research teams achieve a capacity for discovery beyond that of individual models.

CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
CERA-MoA:与持续学习 LLM Agent 协同进化的路由机制
arXiv:2609.18779 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文设计了一种预测式熟悉度估计器,利用中间层隐藏状态评估 Agent 间的语义能力,避免完整 rollout 的开销,并在任务性能和效率之间实现权衡。A predictive familiarity estimator that leverages mid-layer hidden states to evaluate semantic competence among agents, avoiding the overhead of full rollouts and achieving a trade-off between task performance and efficiency is designed.

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
RetireOPD:用于智能体强化学习的自退役 On-Policy 蒸馏
arXiv:2609.20784 Agent 智能体 方法 OA · 绿色 被引 1 · S2

该工作提出 RetireOPD(Self-Retiring On-Policy Distillation),先用环境奖励优化一个解耦的、技能条件化的教师模型,再联合 RL 与 OPD 训练一个无技能依赖的学生模型。This work proposes RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD.

Verifiable Social Reasoning for LLM Assistants
面向 LLM 助手的可验证社会推理
arXiv:2609.17496 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 Fuse——一个用于研究用户中介社会推理的多智能体仿真框架,并将其应用于 12 个 LLM,通过系统性隔离关键因素来展示其分析效用。This work introduces Fuse, a multi-agent simulation framework for studying user-mediated social reasoning, and applies Fuse to 12 LLMs and demonstrates its analytical utility by systematically isolating key factors.

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Designer-RSI:从用户流量中演化过程式记忆用于智能体平面设计
arXiv:2609.22086 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种持续自适应框架:冻结的前沿模型通过超过 230 个工具操作专业设计软件,同时外部的自然语言技能程序性记忆从经验中不断积累并精炼可复用的设计流程。A continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience is introduced.

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
BI-Agent 与 BI-Bench:迈向端到端商业智能自动化
arXiv:2609.20886 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

设计了一个工具增强的 BI-Agent,将 BI 工作流分解为针对结构化数据的子任务(如 search、join、transform),并在各 BI 阶段编排专门的数据管理方法;与此同时开发了一套后训练框架,从真实 BI 项目中合成训练轨迹,使 BI-Agent 能够通过监督微调(SFT)和强化学习(RL)进行进一步后训练。A tool-augmented BI-Agent is designed that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages, and a post-training framework is developed that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL).

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
GAVEL:用于验证且高效的长视野 LLM 任务规划的图世界模型
arXiv:2609.19315 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

改进表明,显式的图世界模型调度层能够显著提升长程具身规划在紧凑型和前沿托管型 LLM 能力下的可靠性和效率。Improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.

Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
arXiv:2609.22255 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Deep Persona,一种心理学驱动的三层架构,将 persona 组织为可观察表达、潜在信念与核心动机驱动的层级结构,用于构建高可信度的角色扮演 Agent。Deep Persona is introduced, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents.

SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness
SkillSpec:面向 Agent Skill 正确性的意图掩码规约推理
arXiv:2609.06052 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出 SkillSpec,一种 Hoare 风格框架,将技能正确性建模为规约推理问题,并将异构技能仓库转化为统一的图表示,对齐描述、指令与代码构件。This work proposes SkillSpec, a Hoare-style framework that formulates skill correctness as a specification reasoning problem, and transforms a heterogeneous skill repository into a unified graph representation that aligns descriptions, instructions and code artifacts.

Lean Pool: An AI-Maintained Archive of Formalized Mathematics
Lean Pool:AI 维护的形式化数学档案库
arXiv:2609.25199 Agent 智能体 方法 OA · 绿色 被引 1 · S2

Lean Pool 是一个形式化数学仓库,由 AI agent 进行生长、维护和优化。Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.

Agensh: Scaling Organizational Intelligence to 1,024 Agents
Agensh:将组织智能扩展到 1,024 个智能体
arXiv:2609.26781 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文揭示 Agent 数量是多 Agent 组织扩展通用智能边界的新 scaling 维度,为硬延迟约束或时间预算下的复杂任务提供了实用方案。The number of agents is revealed as a new scaling dimension for multi-agent organizations to expand the frontier of general intelligence, offering a practical solution for complex tasks under hard latency constraints or time budgets.

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Just-in-Time Memory:面向 LLM Agent 的任务自适应记忆策展学习
arXiv:2609.27334 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Just-in-Time Memory (JitMem) 一致优于无 memory 的 Agent 以及启发式与学习式写入 memory 方法,相较最强基线在成功率上分别提升了 16.2、16.3 和 3.9 个绝对百分点。Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively.

PUBG Ally: A Conversational Embodied Agent as an AI Teammate
PUBG Ally:作为 AI 队友的对话式具身 Agent
arXiv:2609.29837 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 PUBG Ally,一个面向 PUBG: BATTLEGROUNDS 的具身代理,能推理、自主行动并作为语音队友与玩家并肩作战,将代理式工具使用与实时游戏控制相结合。PUBG Ally is introduced, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate that combines agentic tool use with real-time game control.

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
World Action Agent:通过世界动作预演利用 VLM 实现机器人操控
arXiv:2609.29964 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 World Action Agent,一个多代理框架,通过它 VLM 可借助基础工具操控机器人,所有决策均在可视化动作工作空间内完成,在相同骨干下优于端到端 VLA、code-as-policy 代理以及一个可视化框架基线。World Action Agent is presented, a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.

Coding Agents for Generalized Task and Motion Planning Problems
用于广义任务与运动规划问题的编程智能体
arXiv:2609.30233 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文发现编码代理在通用 TAMP 上表现出惊人的有效性:在平均成功率上,三种代理配置均优于人工设计的规划器、一次性生成以及基于 LLM 的通用规划基线。This work finds that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success.

Systems 补充候选
arXiv:2511.02230 Agent 智能体 方法 Open MIND OA · 绿色 被引 68 · S2

Continnum,一种通过为 KV cache 保留引入 TTL 机制来优化多轮 Agent 工作负载任务完成时间的服务系统,与程序级 FCFS 结合时可保持多轮连续性,并降低 Agent 工作流的延迟。Continnum, a serving system to optimize job completion time for multi-turn agent workloads by introducing time-to-live mechanism for KV cache retention, and when combined with program-level first-come-first-serve preserves multi-turn continuity, and reduces delay for agentic workflows.

TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent
TimeEvo:时序 Agent 的失败驱动自演化
arXiv:2609.27277 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

TimeEvo 将 Agent 诊断出的失败聚类为能力缺口,为每个缺口规划测量,合成只产出证据的工具来填补缺口,并仅通过配对的准入门控接纳候选工具库。TimeEvo is proposed, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate.