研究库 论文知识库
Papers · organized/paper_cards

论文

432 张论文卡片 · Agent 智能体

开放获取 全部 绿色 · 1640
Towards In-Parameter Memory Augmentation for Large Language Models
迈向大语言模型的参数内记忆增强
arXiv:2610.08630 Agent 智能体 方法 被引 0 · S2

本综述聚焦在部署阶段为 LLM 增强此类参数化记忆的方法:一个承载记忆的参数对象在推理时被插入到前向传播中,无论该对象是在部署前还是部署中获得。This survey focuses on methods that augment LLMs with such parametric memory at deployment: a memory-bearing parameter object is plugged into the forward pass during inference, whether it is acquired before or during deployment.

MiniCorp: The Last Mile of the AI Agent Firm
MiniCorp:Agent 公司的"最后一公里"
arXiv:2610.05912 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

论文介绍了 MiniCorp,一个用于研究 Agent 如何协作运营公司并规模化生成企业数据的办公模拟器,并对其端到端保真度相对于真实市场实证研究所报告的模式进行了评估。MiniCorp, an office simulator for studying how agents can collectively run a company while generating enterprise data at scale, is introduced and end-to-end fidelity against patterns reported in empirical studies of real markets is evaluated.

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks
DAEDALUS:从自生成任务中引导 Agent 记忆
arXiv:2610.08048 Agent 智能体 方法 被引 0 · S2

本文提出 DAEDALUS,一种无需现有任务或 oracle 验证器、从自生成实践中引导可复用 agent 记忆的方法,并证明在较小探索预算下即可获得性能提升,其启发式策略同样能惠及其他模型家族的 agent。DAEDALUS is presented, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers, and it is shown that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families.

Structuring MoE Expert Selection for Agentic Reinforcement Learning
面向 Agentic 强化学习的 MoE 专家选择结构化
arXiv:2610.07332 Agent 智能体 方法

长程 LLM agent 常借助稀疏 mixture-of-experts(MoE)模型实现,但 agentic 行为与 MoE 结构的协同设计仍缺乏充分探索。本文系统研究 agentic 后训练与 MoE 专家选择之间的关联。在现成 MoE 模型中,我们观察到专家选择呈现出与 agentic 轨迹自然对齐的专门结构:agent 执行语义相似操作(如 READ、UPDATE)的回合之间,其专家路由的重叠度高于执行不同操作……Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing oper

ExperienceIndex: Artifact-Grounded Memory
ExperienceIndex:基于 artifact 的记忆
arXiv:2610.10091 Agent 智能体 方法

知识密集型任务需要通过推理共享的 artifact 语料库(如案件或科学文献)来回答多个问题。随着人类与这些语料库交互,他们自然地积累关于 artifact 的经验知识,从而能够快速识别每个新任务所需的完整相关 artifact 集合。然而,现有 AI agent 缺乏构建或复用此类 artifact-grounded 经验的合适记忆方案,导致答案质量降低且在线成本升高。现有记忆方案从先前任务求解过程中提取并复用信息……(原文截断)Knowledge-intensive tasks require answering many questions by reasoning about a shared corpus of artifacts (e.g., court cases, or scientific literature). As humans interact with these corpora, they naturally accumulate experiential knowledge about artifacts, enabling them to quickly identify the complete set of relevant artifacts for each new task. However, existing AI agents lack appropriate memory solutions to build or reuse such artifact-grounded experience, leading to lower answer quality and higher online cost. Existing memory solutions extract and reuse information from prior task-solvin

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery
从 Pareto 到偏好:通过摊销 agentic 策略发现实现个性化测试时扩展
arXiv:2610.09684 Agent 智能体 方法

测试时扩展(TTS)通过分配额外的推理算力来提升大语言模型的推理能力。现有提升 TTS 效率的方法通常一次针对单一资源维度优化准确率,推进 accuracy–cost 或 accuracy–latency 的 Pareto 前沿。然而用户需求是多维的:用户可能同时指定准确率、延迟与推理成本需求,而不同需求可能偏好不同的控制器。我们将个性化测试时扩展形式化为发现可执行控制器,以最大化……(原文截断)Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
SkillForge:通过动态技能生命周期协同演化技能与 agent
arXiv:2610.09832 Agent 智能体 方法

记忆增强强化学习增强了 LLM agent 解决复杂长程任务的能力。技能是此类记忆的一种形式,将指令与跨任务类型的适用条件配对。然而随着策略改进,若不加区分地保留所有技能,会导致过时或有害条目累积并误导 agent。我们提出 SkillForge,一种 agentic RL 方法,通过由适应性驱动的技能生命周期(试用、活跃、稳定、退役状态)来编译与演化技能库,使技能与模型在整个训练过程中协同演化。预 RL……(原文截断)Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A pre-RL

PhysEvo: Astra Can Act, Let It
PhysEvo:让 Astra 学会行动
arXiv:2610.08995 Agent 智能体 方法

Astra 具备行动能力,但可靠的操作取决于其观察和控制世界的系统。本文提出 PhysEvo,一个围绕单个冻结模型构建的物理递归自我改进(RSI)框架。任务 Agent 执行机器人任务,元 Agent 利用产生的轨迹诊断失败、修订工具与技能,并测试修正效果。元 Agent 还能改进自身的诊断工具,使保留的修订同时支持后续行动与后续自我改进。该过程发展出关节级控制、循证观察以及可复用的操作……Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipula

Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation
通过在线上下文蒸馏将 Agent 经验内化至扩散模型权重
arXiv:2610.07250 Agent 智能体 方法

将图像生成模型包裹在 Agent 框架中可有效提升 Text-to-Image 任务性能:该框架能利用记忆、技能、工作流编排、结果验证与迭代优化来持续构建和修订 prompt,从而生成更好的图像。然而这些增益对扩散模型而言是外部的,只有运行完整框架时才能实现。本文提出扩散在线上下文蒸馏(D-OPCD),将 Agent 改进后的 prompt 视为特权上下文,并将编码在 Agent 框架中的知识蒸馏进……Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby eliciting better images. These gains, however, remain external to the diffusion model and are realized only while the full harness runs. We propose Diffusion On-Policy Context Distillation (D-OPCD), which treats the agent-improved prompt as privileged context and distills the knowledge encoded in the agent harness into the

UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
UniSkill:为持续进化的策略学习与执行者对齐的技能提案
arXiv:2610.10164 Agent 智能体 方法

大语言模型 Agent 可通过保留从先前交互中提炼的可复用技能来跨任务提升能力。已有研究联合优化任务执行与技能提取,使策略与技能库协同进化。然而,随着执行者持续学习,通过技能在后续训练步骤中的复用对其进行奖励,可能将技能收益与执行者自身的改进混淆;而直接测试每个候选技能又需要代价高昂的额外执行者 rollout。本文提出 UniSkill,利用共享策略与环境交互并提出技能……Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions. Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve. However, as the actor continues learning, rewarding skill proposals through their reuse in subsequent training steps may conflate skill benefits with actor improvement, while directly testing each proposed skill requires costly additional actor rollouts. In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose ski

A self-learning scientific agent for X-ray diffraction
用于 X 射线衍射的自学习科学 Agent
arXiv:2610.07862 Agent 智能体 观点

科学 Agent 的核心挑战在于将分析经验转化为基于物理证据的可复用专业知识。本文介绍 Gan Jiang,一个基于我们自主开发的衍射分析生态系统(XMatcher、XQueryer、XDecomposer 与 WPEM)构建的面向粉末 X 射线衍射的自学习 Agent。这些引擎共同涵盖物相鉴定、多相分解以及物理约束的全谱建模。Gan Jiang 通过诊断失败、修订技能指令与代码,并在复用前验证修订,将分析经验转化为可执行的技能,伴随……A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by diagnosing failures, revising skill instructions and code, and validating revisions before reuse, wi

CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use
CADFather:通过协同工具调用实现自主 CAD 重建
arXiv:2610.09127 Agent 智能体 方法

从三维形状重建可编辑 CAD 模型仍是一项具有挑战性的工程任务。现有方法可以提出 CAD 操作,但没有任何单一提案源能对不同零件几何与不同重建阶段同样有效。本文介绍 CADFather,一个自主 Agent 系统,通过协调互补工具从三维网格恢复参数化 CAD 程序。视觉-语言助手检查目标与中间重建结果的渲染图像,再决定扩展哪些候选 CAD 程序、调用哪些工具、生成多少提案,以及……Reconstructing an editable CAD model from a 3D shape remains a challenging engineering task. Existing methods can propose CAD operations, but no single source of proposals works equally well across different part geometries and stages of reconstruction. We introduce CADFather, an autonomous agentic system that coordinates complementary tools to recover parametric CAD programs from 3D meshes. A vision-language assistant inspects renders of the target and intermediate reconstructions, then decides which candidate CAD programs to extend, which tools to invoke, how many proposals to generate, and

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
自回顾蒸馏:将事后经验转化为先验远见
arXiv:2610.08077 Agent 智能体 方法

可验证奖励的强化学习 (RLVR) 主要通过交互后的标量结果奖励将 Agent 经验转化为学习信号。然而对于组相对目标,当所有推演获得相同奖励时,该信号即消失,尽管这些轨迹可能包含关于任务内容及 Agent 失败方式的有用信息。我们提出一个互补问题:事后回顾能否教会 Agent 在行动前本可预见的内容?我们引入前瞻学习,利用事后经验从行动前的Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-

Agent Plasticity: Measuring Self-Improvement Through Experience
Agent 可塑性:通过经验衡量自我改进
arXiv:2610.08902 Agent 智能体 方法

AI Agent 日益在能够诊断失败并通过经验改进的环境中运行,然而现有评估主要衡量 Agent 在固定时间点能做什么,而非其学习效果如何。评估自我改进需要回答三个问题:未来性能是否改进并泛化至学习交互之外;新能力获取效率如何;自我改进过程在哪里失效?为回答这些问题,我们在受控环境中研究自我改进,其中 Agent 分摊AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future performance improve and generalize beyond the interactions that enabled learning; how efficiently are new capabilities acquired; and where does the self-improvement process break down? To answer these questions, we study self-improvement in a controlled setting where agents amortize pa

Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance
Inherit-MAS:通过工作流与执行继承实现多 Agent 系统的测试时演化
arXiv:2610.02396 Agent 智能体 方法

由大语言模型构建的多 Agent 系统 (MAS) 通过协调专业化 Agent 来处理复杂任务,但有效的工作流难以预先设计。测试时演化利用执行反馈来优化工作流,然而广泛的修订可能扰动有用组件,而重执行未变更的请求会带来冗余计算。受生物进化中继承与选择相互作用的启发,我们引入 Inherit-MAS,在工作流和执行层面显式化继承。元模型首先综合出由 worker Agent 组成的工作流Multi-agent systems (MAS) built from large language models coordinate specialized agents to tackle complex tasks, but effective workflows are difficult to design in advance. Test-time evolution refines workflows using execution feedback, yet broad revisions can disturb useful components, while re-executing unchanged requests can incur redundant computation. Inspired by the interplay of inheritance and selection in biological evolution, we introduce Inherit-MAS, which makes inheritance explicit at the workflow and execution levels. A meta-model first synthesizes a workflow of worker agents with

Forms of LLM-Integrated Applications from LLM-Chats to Autonomous AI Agent System
从 LLM-Chat 到自主 AI Agent 系统中 LLM 集成应用的形态
arXiv:2610.11899 Agent 智能体 综述

LLM 越来越多地作为组件嵌入软件系统,并以 chatbot、copilot、RAG、workflow、coding agent、AI agent 等标签推广。这些标签究竟指代真实的架构形态,还是仅作品牌包装,尚未得到系统性评估。在调研的来源中,标签确实承载架构含义,在厂商用法中体现得最为清晰:copilot 指在逐步用户确认下操控宿主应用的 router-worker 架构,而近来向 agent 标签的迁移则与 AI 自主规划相吻合……Large language models (LLMs) are increasingly embedded as components in software systems, marketed under labels such as chatbot, copilot, retrieval-augmented generation, workflow, coding agent and AI agent. Whether these labels denote genuine architectural forms or serve as branding has not been assessed systematically. In the sources surveyed, labels do carry architectural content, most clearly in vendor usage: copilot denotes a router-worker architecture operating a host application under step-by-step user confirmation, while the more recent shift to the label agent coincides with AI-planned

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
Memento 3:通过反思式规则手册实现基于模型的递归自我改进
arXiv:2610.11794 Agent 智能体 方法

在未知环境中学会行动需要智能体推断世界运作方式,并在新证据出现时修正理解。然而,有限的观察可能支持多个世界模型,它们都能解释过去的交互,但对未见状态有不同的预测。我们提出 Memento 3,在 Memento 系列的基础上,使冻结的 LLM 智能体能够通过外部记忆持续学习显式世界模型。智能体维护一个自然语言规则手册作为持久化语义记忆,记录可修正的环境动态假设,同时保留未知部分Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects

A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
深入探究 Agentic BBO:面向黑盒优化的 LLM 智能体基准测试
arXiv:2610.12183 Agent 智能体 评测集

黑盒优化(BBO)出现在许多目标和工程问题中,其目标函数评估代价高昂且次数有限。近期的大语言模型(LLM)智能体通过结合任务语义、计算、优化工具以及反馈驱动的决策,提供了一种新的 BBO 解决方式,因与数学严谨工具的集成而展现出巨大潜力。然而,现有的 Agentic BBO 研究使用不同的任务领域和系统配置,导致结果难以比较,且单个设计选择的影响难以孤立分析。因此,我们引入Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools. However, existing agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We therefore i

REMORY: Learning Residual Memory for Context Compaction
REMORY:学习用于上下文压缩的残差记忆
arXiv:2610.11287 Agent 智能体 方法

长时程智能体压缩其历史以在有限上下文窗口内继续执行,但单一的文本摘要可能无法支撑所有后续决策。我们提出 REMORY,一种神经记忆网络,通过有界序列的软记忆 token 来补充摘要。给定历史和摘要,该网络学习生成帮助冻结 LLM 近似其在完整历史下产生的续写内容的 token。这些 token 以摘要为条件并附加在其后,构成沿序列维度的残差连接类比。在 SummHay 上,REMORY 提升了Long-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the history and summary, the network learns to generate tokens that help a frozen LLM approximate the continuation it would produce with the full history. The tokens are conditioned on the summary and appended after it, forming an analogue of a residual connection along the sequence dimension. On SummHay, REMORY im

Opera: A Verbal Critic Framework for Long-horizon Coding Agents
arXiv:2609.33987 Agent 智能体 方法

Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evi

Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub
arXiv:2610.11169 Agent 智能体 方法

Agent skills are SKILL.md instructions and scripts that AI coding agents such as Claude Code and Codex run with the permissions of their user. Developers share skills by copying them between repositories, which makes them a software supply chain without a registry, versions or provenance. The origin of a copied skill, the reach of a security fix and the repositories that warrant review are therefore unknown. Studies that record which repositories hold a skill at a single point in time cannot reveal who copied it from whom. We contribute the first dated copy network of agent skills, built from

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
CanvasAgent:通过视觉工具编排实现复杂图像创建与编辑
arXiv:2607.05465 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出了用于复杂图像创建与编辑的大规模多模态工具调用数据集 CanvasCraft,以及通过多轮交互学习编排异构视觉工具的工具增强多模态 Agent CanvasAgent。CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and CanvasAgent, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction are introduced.

Multi-Turn Agentic Scientific Literature Search via Workflow Induction
基于工作流归纳的多轮智能体科学文献搜索
arXiv:2607.00597 Agent 智能体 方法 OA · 绿色 被引 2 · S2

结果表明,显式且可编辑的搜索工作流为将文献搜索智能体与复杂科学意图对齐提供了有效且可控的接口。The results show that explicit, editable search workflows provide an effective and controllable interface for aligning literature search agents with complex scientific intent.

Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Care
教导 LLM 在欠发达的癫痫诊疗中进行推荐与转诊
arXiv:2606.31036 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

在资源受限环境中,专业癫痫专家稀缺,使基于 LLM 的决策支持对管理纵向治疗的一线临床医生具有吸引力。此类系统必须适应当地处方实践并知道何时转诊。我们在乌干达儿科癫痫诊疗中研究该问题,基于纵向非结构化门诊记录预测抗癫痫用药方案。标准提示与医生处方取得了一定程度的一致性,但神经科医生审查显示许多错误反映的是分布失校的处方默认值而非失败。Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractive for frontline clinicians managing longitudinal treatment. Such systems must adapt to local prescribing practice and know when to defer. We study this problem in Ugandan pediatric epilepsy care, predicting anti-seizure medication regimens from longitudinal unstructured clinic notes. Standard prompting achieves non-trivial agreement with physician prescriptions, but neurologist review shows that many errors reflect distribution-miscalibrated prescribing defaults rather than fail

A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory
A-TMA:解耦长时 Agent 记忆中的状态感知失效
arXiv:2607.01935 Agent 智能体 方法 OA · 绿色 被引 4 · S2

提出 ATMA:在现有记忆系统之上的状态感知叠加层,保留被替换记录与过渡记录,为查询所需的"目标状态视图"构建证据包,并向问答模块暴露当前、历史与过渡三类标签。This work proposes ATMA, a state aware overlay for existing memory systems, which keeps superseded and transition records in the bank, builds evidence packets for the query's requested state view, and exposes current, historical, and transition labels to QA.

AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
AgenticSTS:面向长时 LLM Agent 的有界记忆测试平台
arXiv:2607.02255 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出了一种 Agent 设计和一套经过验证、可复用的方法,用于研究显式记忆层如何影响长周期 LLM Agent 决策,并给出了一种替代的有界契约。An agent design and a validated, reusable methodology for studying how explicit memory layers shape long-horizon LLM-agent decisions, as well as an alternative bounded contract, are introduced.

TRACE: State-Aware Query Processing over Temporal Evidence Graphs for Conversational Data
TRACE:面向会话数据的时序证据图上的状态感知查询处理
arXiv:2607.00339 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 TRACE,一个针对演化会话数据在时序证据图上的查询处理框架,将词汇召回与证据重建分离,从而在长会话历史中实现有界的查询时推理。TRACE is presented, a query processing framework over temporal evidence graphs for evolving conversational data that separates lexical recall from evidence reconstruction, enabling bounded query-time reasoning over long conversational histories.

Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
将个性化建模为逆向规划:通过结构去噪学习潜在设计意图以实现 Agentic 幻灯片生成
arXiv:2607.00407 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文将 PSP 形式化为逆向规划问题,并提出 SPIRE,一个通过有意破坏干净幻灯片的视觉结构来近似求解 PSP 的原则性框架,从而构建一个可验证的去噪任务。This work formulates PSP as an inverse planning problem, and proposes SPIRE, a principled framework to solve PSP approximately, by intentionally corrupting the visual structures of clean slides, which creates a verifiable task to denoise the corruption.

A Large-Language-Model Supported Personalized Driving Framework for Lane Change in Highway Scenarios
一种由大语言模型支持的面向高速公路场景换道的个性化驾驶框架
arXiv:2606.31483 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

实验结果表明,所得参数集可生成可区分的个性化换道行为,同时 RAG 始终提升偏好理解效果,尤其对隐式指令效果显著,表明将基于 LLM 的自然语言交互与 Apollo 集成以支持个性化换道行为生成具有潜力。Experimental results show that the derived parameter sets generate distinguishable personalized lane-change behaviors, while RAG consistently improves preference interpretation, particularly for implicit commands, indicating the potential of integrating LLM-based natural-language interaction with Apollo to support personalized lane-change behavior generation.

Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming
保障 AI Agent 安全:面向多层 Agent 红队测试的统一框架
arXiv:2606.31227 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本文提出 AI-Infra-Guard,一个围绕单一观测组织 AI 红队测试的开源框架,是目前唯一覆盖所有层面(包括对日益扩展 AI Agent 能力的 Agent Skills 供应链审计)的开源框架。AI-Infra-Guard is presented, an open-source framework that organizes AI red teaming around a single observation, and is the only open-source framework to span all of these, including supply-chain auditing of the agent skills that increasingly extend AI agents.

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking
PixelEyes:解耦感知与推理以实现精准视觉证据定位
arXiv:2607.00115 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文探索多轮视觉推理,观察到 MLLM 反复无法定位目标,导致冗长的推理轨迹,进而提出 PixelEyes,一种将推理与感知显式解耦的多轮视觉推理 Agent。This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories, and proposes PixelEyes, a multi-turn visual reasoning agent that explicitly decouples reasoning from perception.

DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation
DuoMem:基于双空间蒸馏的可行端侧记忆 Agent
arXiv:2606.29961 Agent 智能体 应用落地 OA · 绿色 被引 1 · S2

本文提出 DuoMem,一种双空间蒸馏框架,可将程序化问题求解能力从大型教师模型迁移到紧凑学生模型,并适用于实时边缘部署,而这一点对教师模型而言颇具挑战。DuoMem is introduced, a dual-space distillation framework that transfers procedural problem-solving ability from a large teacher model to compact student models and is viable for real-time edge deployment, which would be challenging for the teacher.

Agentic Abstention: Do Agents Know When to Stop Instead of Act?
Agentic Abstention:Agent 是否知道何时该停止而非行动?
arXiv:2606.28733 Agent 智能体 方法 OA · 绿色 被引 13 · S2

研究发现,模型规模、推理能力与 Agent 脚手架以不同方式影响弃答行为,能力更强或更大的模型有时反而在及时弃答上表现更差。It is found that model scale, reasoning, and agent scaffolding affect abstention in different ways, where larger or more capable models sometimes perform worse at timely abstention.

When Multi-Robot Systems Meet Agentic AI:Towards Embodied Collective Intelligence
当多机器人系统遇上 Agentic AI:迈向具身集体智能
arXiv:2606.27929 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文探讨了具身集体智能 (ECI) 这一未来多机器人范式,其中机器人团队将世界上下文、任务进度与技能经验作为共享资源进行累积与利用。This article explores Embodied Collective Intelligence (ECI), a future multi-robot paradigm in which a robot team accumulates and uses world context, task progress, and skill experience as shared resources.

From Detection to Action: Using LLM Agents for Fault-Tolerant Control
从检测到行动:基于 LLM Agent 的容错控制
arXiv:2606.28011 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文提出一种基于 Agentic Large Language Model (LLM) 的主动容错控制 (FTC) 框架,可将故障检测输出转化为基于特定工厂知识的、符合约束的恢复动作。该方法结合:(i) 将操作员职责分解为监测、规划、动作合成、仿真、验证与重新提示的多 Agent 工作流;(ii) 数字过程工厂孪生 (DPPT),提供工厂数据、模型以及用于执行前测试的仿真服务;(iii) 基于 CPSMod 本体构建的 Graph Retrieval-Augmented Generation (Graph RAG) 层。We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constraint-aware recovery actions grounded in plant-specific knowledge. The approach couples (i) a multi-agent workflow that decomposes operator duties into monitoring, planning, action synthesis, simulation, validation, and reprompting; (ii) a Digital Process Plant Twin (DPPT) that exposes plant data, models, and a simulation service for pre-execution testing; and (iii) a Graph Retrieval-Augmented Generation (Graph RAG) layer built on the CPSMod ontol

GBC: Gradient-Based Connections for Optimizing Multi-Agent Systems
GBC:用于多智能体系统优化的基于梯度的连接
arXiv:2606.28187 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Gradient-Based Connections(GBC)——一种面向多智能体系统的细粒度归因与优化方法,可提升多智能体性能,超越强力的单智能体与多智能体基线;且归因质量越高,优化效果越显著。Gradient-Based Connections (GBC) is proposed, an approach for fine-grained attribution and optimization of multi-agent systems that improves multi-agent performance and outperforms strong single-agent and multi-agent baselines and higher attribution quality is associated with greater optimization effectiveness.