提出 Single-rollout Asynchronous Optimization(SAO),用于解决异步 RL 中的稳定性与 off-policy 难题,可稳定训练一千步,并在 Agentic 编码与推理基准上一致优于 GRPO 及其变体。Single-rollout Asynchronous Optimization (SAO) is presented to address the stability and off-policy challenges in asynchronous RL and is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks.
论文
432 张论文卡片 · Agent 智能体
提出 AgentLens,一个面向交互式代码 Agent 的生产级评估基准,将形式化验证(在存在客观检查时)与 LLM 编写的轨迹评审及并排比较相结合,使每次运行都能给出关于分数为何如此的可读解释。This work presents AgentLens, a production-assessed benchmark for interactive code agents that pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is.
提出首个面向动态真实世界场景评估主动 agent 的能力驱动基准,并在模型与框架层面的全面比较表明,基模型能力与 agent 框架设计共同决定了真实世界环境中的性能表现。The first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world settings is introduced, and comprehensive comparisons across both models and frameworks show how base model capabilities and agent framework designs jointly shape performance in real-world environments.
提出 TokenWall,一种作用于 agent token 流的语义防火墙式运行时防御框架,证明语义运行时约束可在持久化 AI agents 上实现实用的安全性与效用性权衡。TokenWall is proposed, a runtime defense framework that acts as a semantic firewall over agent token flows, demonstrating that semantic runtime containment can achieve a practical security-utility trade-off for persistent AI agents.
提出 Contextuality(上下文性)——即 AI 系统自主访问用户累积知识资本的程度——作为 AI 介导不平等的一个维度,补充但不可化约为 Sharp 等人的框架。Contextuality -- the degree to which an AI system autonomously accesses a user's accumulated knowledge capital -- is proposed as a dimension of AI-mediated inequality that complements, but is not reducible to, the Sharp et al. framework.
消融实验表明,选择性干预优于被动记忆库暴露、常驻注入、仅顾问引导和通用检索。Ablations show that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, advisor-only guidance, and general retrieval, and general retrieval and that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, and general retrieval.
该工作提出 LongMedBench,一个基于真实 EHR 的长程临床决策基准,并设计了一套包含三个评估维度的分类体系:事实型问答、时序推理、长程决策。This work introduces LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making, and proposes an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making.
该工作提出 Long-Horizon-Terminal-Bench,一个涵盖九大类共 46 个长程任务的终端基准,包括实验复现、软件工程、多模态分析、交互游戏与科学计算,并分析失败模式与错误规律,以推动长程终端 Agent 的后续研究。This work introduces Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing, and analyzes failure modes and error patterns to support future progress on long-horizon terminal agents.
论文提出 Flow-ERD,一个同时追求真实性与多样性的多 Agent 仿真器,在 WOSAC 测试基准上排名第一,并在可复现基线中主导真实性-多样性 Pareto 前沿。Flow-ERD is introduced, a multi-agent simulator that pursues realism and diversity jointly and ranks first on the WOSAC test benchmark and dominates the realism--diversity Pareto front among reproducible baselines.
本文首次全面综述了 LLM 元认知的研究现状,涵盖用于测量和评估 LLM 元认知能力的方法与 benchmark、激发、改进与应用 LLM 元认知的技术,以及当前研究的发现与启示。The first comprehensive overview of the current state of knowledge on metacognition for LLMs is presented, including methods and benchmarks to measure and evaluate LLMs'metacognitive abilities, techniques to elicit, improve, and apply metacognition in LLMs, and findings and implications of ongoing research.
提出 ABot-AgentOS,一个通用机器人 Agent Operating System,位于底层控制器之上,提供 deliberation agent 层,支持场景条件规划、上下文隔离的 Skill 执行、多阶段验证、多模态记忆以及边云协同。ABot-AgentOS is presented, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration.
本文提出 Multi-Agent Contextual Exploration (MACE),一个通过结构化的对等体选择显式促进探索的轻量级框架,显著改善了探索行为和下游任务表现,并在理论上证明探索价值随 Agent 多样性增加而提升。This work introduces Multi- Agent Contextual Exploration (MACE), a lightweight framework that explicitly promotes exploration through structured peer selection that substantially improves exploration behavior and downstream task performance and shows theoretically that the value of exploration increases with agent diversity.
基于 LLM 的编程 Agent 显著推动了自动化软件问题解决,但由于对仓库理解不足,仍易出现事实性错误。近期方法尝试通过修复前仓库探索来缓解此问题;然而,其修复驱动策略在未识别 Agent 知识缺口的情况下探索仓库,往往产生不精确的上下文,无法弥补潜在的理解不足。本文提出 ACQUIRE,一种面向软件问题解决的 QA 驱动框架,模拟经验丰富的开发者LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods attempt to mitigate this limitation through pre-repair repository exploration; however, their fix-driven strategies explore repositories without identifying the agent's knowledge gaps, often yielding imprecise context that fails to bridge the underlying understanding deficit. In this paper, we propose ACQUIRE, a QA-driven framework for software issue resolution. Mirroring how experienced developer
介绍 AMID,一种面向医学影像模型开发的自主多 Agent 框架,其性能优于所评估的通用 MLE 系统,并在异构任务上接近或匹配强大的人工设计挑战赛方案。AMID is introduced, an autonomous multi-agent framework for medical imaging model development that outperformed evaluated general-purpose MLE systems and approached or matched strong human-designed challenge solutions across heterogeneous tasks.
除领域内增益外,mid-training 还能缓解 agentic post-training 对非 Agent 编程及非编程工具调用基准(tau-bench、BFCL)造成的能力侵蚀:尽管 mid-training 语料仅含 Python 代码,函数调用的归纳偏置在 post-training 后依然保留,带来稳定的增益。Beyond in-domain gains, mid-training mitigates the capability erosion that agentic post-training otherwise inflicts on non-agent coding and non-coding tool-use benchmarks (tau-bench, BFCL): although the mid-training corpus contains Python code only, the function-call inductive bias survives post-training and yields consistent gains.
尽管视觉-语言模型(VLMs)已取得成功,但误导性图表因欺骗性视觉结构与失真数据表示仍构成重大挑战。我们提出 ChartCynics,一个通过"怀疑式"推理范式揭露视觉欺骗的 Agentic 双路径框架。与整体化模型不同,ChartCynics 将感知与验证解耦:诊断式视觉路径通过策略性 ROI 裁剪捕获结构异常(如倒置坐标轴),OCR 驱动数据路径确保数值根植性。为解决跨模态冲突,我们提出Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework designed to unmask visual deception via a "skeptical" reasoning paradigm. Unlike holistic models, ChartCynics decouples perception from verification: a Diagnostic Vision Path captures structural anomalies (e.g., inverted axes) through strategic ROI cropping, while an OCR-Driven Data Path ensures numerical grounding. To resolve cross-modal conflicts, we introduce
本研究将朴素搜索的根因追溯到生成器特有的、可演化的知识边界——即生成器经训练可内化的内容与必须保留于外部上下文的内容之间的鸿沟,并表明该边界可通过"先教后搜"协同训练框架被有效发现。This work traces the root cause of naive search to a generator-specific, evolving knowledge boundary: the divide between what a generator can internalize through training and what must remain in external context, and shows that it is discoverable through a teach-then-search co-training framework.
本研究考察了 GitHub 项目在各自引入首个 bot 前后两年间的情况,发现变化集中在采纳时点附近,而非逐渐累积,这与一种特定解读一致:可预测、基于规则的 Agent 能够成为社区社交基础设施的一部分。This work examines GitHub projects for two years before and after each adopted its first bot, finding changes cluster around adoption rather than accumulating gradually, consistent with a specific interpretation: predictable, rule-based agents can become part of a community's social infrastructure.
论文提出 Vinci2,一个主动式的第一人称视频协助系统,将端侧助手 Vinci 由被动响应推进到主动协助;以及免训练、记忆增强的 Agent EgoMemo,维护三种互补的记忆表征:多尺度时间摘要、语义知识图谱与视觉嵌入档案。Vinci2 is presented, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity and EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives.
论文提出 OAT,将该问题建模为基于神经受控微分方程的单类学习,在潜空间中刻画成功轨迹的动力学模式;实验表明其比基于 prompt 的基线更快,并在领域内和分布外数据集上均稳定优于基线。OAT is proposed, which casts this problem as one-class learning with neural controlled differential equations, modeling the dynamical pattern of successful trajectories in latent space, and is shown to be faster than prompting-based baselines and consistently outperforms them in both in-domain and out-of-distribution datasets.
STRACE(Structural TRajectory Analysis and Causal Extraction)是一个用于构建高信噪比优化上下文的框架,旨在对长周期 Agent 实施更精确、更有效的优化。STRACE (Structural TRajectory Analysis and Causal Extraction) is a framework that constructs high signal-noise optimization contexts for more precise and effective optimization of long-horizon agents.
PalmClaw 是一个开源 Agent 框架,原生运行于手机端,直接在设备上管理 session、memory、Skill、工具以及 agent loop,使 Agent 能够直接调用移动端能力,同时保证每一步操作的显式与可控。PalmClaw is an open-source agent framework that runs natively on mobile phones and manages the sessions, memory, skills, tools, and agent loop directly on the device, allowing agents to use mobile capabilities directly while keeping each action explicit and controlled.
本综述将现代具备自我改进能力的 Agent 视为将经验转化为持续能力增益的自适应系统,并提出一个系统级框架,将现代 Agent 建模为由基础模型与由 prompt、memory、工具及控制逻辑构成的运行支撑层相耦合的配置。This survey frames modern self-improving agents as adaptive systems that convert experience into accumulated capability gains, and offers a system-level framework that represents a modern agent as a configuration coupling a foundation model with an operational scaffold of prompts, memory, tools, and control logic.
SPEAR 是一个 Python 库,可通过模块化插件架构连接任意 Unreal Engine 应用并对其进行编程化控制;同时引入一种表达力强的高层编程模型,使用户能够以任意数据依赖关系指定复杂的 UE 工作图,并在单个 UE 帧内确定性执行这些图。SPEAR is a Python library that can connect to, and programmatically control, any Unreal Engine application via a modular plugin architecture, and introduces an expressive high-level programming model that enables users to specify complex graphs of UE work with arbitrary data dependencies among work items, and to execute these graphs deterministically within a single UE frame.
提出一个多 Agent 框架,通过结合监督微调、直接偏好优化和检索增强生成 (RAG) 来调和事实基础与意识形态对齐,产生稳定的胜者和排名,且以宣言为锚的谱系能可靠预测现实世界中的实现,而幻觉内容则不能。A multi-agent framework that reconciles factual grounding with ideological alignment by combining Supervised Fine-Tuning, Direct Preference Optimization, and Retrieval-Augmented Generation is presented, which yields a stable winner and ranking, and manifesto-anchored lineage reliably predicts real-world materialization whereas hallucinated content does not.
提出 Hy-Embodied-RxBrain,一个具备语言-视觉联合推理与想象的具身认知基础模型,并将其扩展到连续机器人动作生成,在无需大规模动作数据预训练的情况下展现出可观的真实机器人性能。Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination, is introduced and extended to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining.
Reflexion 是一个通过语言反馈而非更新权重来强化语言 agent 的新框架,在多种任务(序贯决策、编程、语言推理)上相较基线 agent 取得显著提升。Reflexion is a novel framework to reinforce language agents not by updating weights, but instead through linguistic feedback, which obtains significant improvements over a baseline agent across diverse tasks (sequential decision-making, coding, language reasoning).
HuggingGPT 是一个由 LLM 驱动的 Agent,利用 LLM(如 ChatGPT)连接机器学习社区中的各种 AI 模型以解决 AI 任务,能够处理跨模态、跨领域的大量复杂 AI 任务。HuggingGPT is an LLM-powered agent that leverages LLMs (e.g., ChatGPT) to connect various AI models in machine learning communities to solve AI tasks and can tackle a wide range of sophisticated AI tasks spanning different modalities and domains.
一篇关于基于 LLM 的 Agent 的全面综述,追溯了 Agent 概念从其哲学起源到在 AI 中的发展历程,解释了为何 LLM 适合作为 Agent 的基础,并提出一个包含三个核心组件的通用框架:大脑、感知与行动。A comprehensive survey on LLM-based agents, tracing the concept of agents from its philosophical origins to its development in AI, and explaining why LLMs are suitable foundations for agents, and presenting a general framework, comprising three main components: brain, perception, and action.
本文在交通微观仿真器 SUMO 中应用现代深度强化学习方法构建一个真正自适应的交通信号控制智能体,并采用一种新的状态空间——离散交通状态编码——其信息密度较高。This work applies modern deep reinforcement learning methods to build a truly adaptive traffic signal control agent in the traffic microsimulator SUMO, using a new state space, the discrete traffic state encoding, which is information dense.
提出 Lucid,一个黑盒对抗框架,在严格的图像受限威胁模型下攻击多模态记忆管道,无需访问目标 MLLM、目标检索编码器或文本通道,揭示了多模态记忆管道中的结构性漏洞。Lucid is proposed, a black-box adversarial framework that compromises multimodal memory pipelines under a strictly image-bounded threat model, requiring no access to the target MLLM, target retrieval encoder, or the text channel, exposing a structural vulnerability in multimodal memory pipelines.
提出数据科学世界模型概念,通过基于当前工作流状态和候选操作预测环境状态转移来建模数据科学执行环境;提出DSWorld框架,结合结构化状态构建、成本感知路由、轻量级真实执行以及基于LLM的昂贵操作模拟器。The concept of Data Science World Model is introduced, which model the data science execution environment by predicting environment state transitions conditioned on current workflow states and candidate operations and proposes DSWorld, a practical framework that combines structured state construction, cost-aware routing, lightweight real execution, and an LLM-based simulator for expensive operations.
TARS是一个集成在Visual Studio Code中的LLM驱动Agent,通过直接锚定到被分析代码的自主解释来支持程序理解,基于轻量级心智理论范式构建。TARS is an LLM-powered agent integrated into Visual Studio Code that supports program comprehension through autonomous explanations anchored directly to the code under analysis, built around a lightweight Theory of Mind paradigm.
在一个 recipe 级操作机制中,fan-in Muon 在共享 KL 与 clipping 下支持更激进的稳定有效步长:该余量在优化仍有空间时最大,而在接近饱和、经 AdamW 调参后或使用 magnitude matching 时收缩。A recipe-level operating regime in which fan-in Muon supports a more aggressive stable effective step under shared KL and clipping is identified: the margin is largest when optimization headroom remains and contracts near saturation, after AdamW tuning, or under magnitude matching.
提出RESOURCE2SKILL框架,将教程视频、仓库、文章和参考制品等多模态资源蒸馏为软件Agent的可执行技能,并验证了多模态技能格式、层次化组织、来源多样性、选择策略与在线获取的价值。RESOURCE2SKILL is presented, a framework that distills multimodal resources, including tutorial videos, repositories, articles, and reference artifacts, into executable skills for software agents, and confirms the value of multimodal skill format, hierarchical organization, source diversity, selection strategy, and online acquisition.
一种面向语言模型推理的新框架 Tree of Thoughts (ToT),推广了流行的 Chain of Thought 提示方法,允许在作为问题求解中间步骤的连贯文本单元(thoughts)上进行探索。A new framework for language model inference, Tree of Thoughts (ToT), which generalizes over the popular Chain of Thought approach to prompting language models, and enables exploration over coherent units of text (thoughts) that serve as intermediate steps toward problem solving.