研究库 论文知识库
Papers · organized/paper_cards

论文

14 张论文卡片 · Agent 智能体 · 观点 · OA 绿色

开放获取 全部 绿色 · 1640
5. Context-Fractured Decomposition Attacks on Tool-Using LLM Agents
5. 上下文碎裂分解攻击针对使用工具的 LLM Agent
arXiv:2606.09084 Agent 智能体 观点 OA · 绿色 被引 3 · S2

揭示使用工具的 LLM Agent 的一种部署失效模式——来源缺口,以及一类跨上下文多步越狱攻击,可在早期交互中保留看似无害的中间产物,并在很久以后(可能在不同 Agent 实例或工作流阶段)诱发有害行为。A deployment failure mode for tool-using LLM agents, the provenance gap, and a family of cross-context multi-step jailbreaks that preserve benign-looking intermediate artifacts from an early interaction and elicit harmful behavior much later, potentially in a different agent instance or workflow stage.

A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
一种面向 CT 扫描中空间关系验证的、可靠且可审计的模块化 Agent
arXiv:2608.21140 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个模块化的医学影像 Agent,用于轴位 CT 切片中的二元空间关系验证,采用显式模块化空间验证阶段,并表明显式模块化空间验证可作为未来面向报告的医学影像 Agent 的有前景的构建模块。This work presents a modular medical imaging agent for binary spatial relation verification in axial CT slices using explicit modular spatial verification stages, and suggests that explicit modular spatial verification can serve as a promising building block for future report-oriented medical imaging agents.

Procedura: Agentic 3D Modeling with Procedural Control
Procedura:基于过程化控制的 Agentic 三维建模
arXiv:2608.26238 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文探索"以代码表示 3D 形状"的范式,利用并放大 LLM 的编码能力进行 3D 建模,并提出 Procedura 框架——通过编写一个由命名零件构成、并通过类型化、可机器校验的连接关系装配而成的参数化程序来建模对象。The paradigm of 3D shape as code is explored, leveraging and scaling the coding ability of an LLM for 3D modeling, and Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates.

Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
编织视觉叙事:超越原子视觉匹配的 Agentic Image Bundle Composition
arXiv:2608.28695 Agent 智能体 观点 OA · 绿色 被引 1 · S2

大量实验表明,SOTA embedding 模型和静态 decompose-and-rerank 范式存在关系盲区,而 BundleWeaver 取得了显著性能提升,凸显了从原子打分转向动态关系组合的必要性。Extensive experiments demonstrate that while state-of-the-art embedding models and static decompose-and-rerank paradigms suffer from relational blindness, BundleWeaver achieves substantial performance gains, highlighting the necessity of shifting from atomic scoring to dynamic relational composition.

4️⃣ arXiv · Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use(⭐⭐⭐⭐ 高优先级)
4️⃣ arXiv · 守护 Agent:厂商中立的多租户企业级检索与工具调用(⭐⭐⭐⭐ 高优先级)
arXiv:2605.05287 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文提出一种分层隔离架构,结合策略感知的 ingestion、retrieval-time gating 与共享推理,并通过服务端 agentic 编排加以执行,在为多租户隔离提供天然强制点的同时,允许客户端框架保留对 agent 组合与延迟敏感操作的控制权。A layered isolation architecture combining policy-aware ingestion, retrieval-time gating, and shared inference, enforced through server-side agentic orchestration is introduced, creating natural enforcement points for multitenant isolation while allowing client-side frameworks to retain control over agent composition and latency-sensitive operations.

Substrate-Aware AI Agents: Execution Context as a First-Class Input
Substrate-Aware AI Agent:将执行上下文作为一等输入。
arXiv:2609.05232 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

最小执行契约可在生成程序中引发主动的结构适应,将计算从无约束分配中移开,在执行前显著改善观察到的资源时间分布,建立了 substrate-aware Agent 规划的受控概念验证。A minimal execution contract induces proactive structural adaptation in generated programs, shifting computation away from unconstrained allocations and substantially improving observed resource-time profiles before execution, establishing a controlled proof of concept for substrate-aware agent planning.

Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
双层协同反思:多 Agent LLM 系统的博弈论方法
arXiv:2609.02750 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出随机反思记忆提升(SRMA),仅在 grounding 后的评估风险严格下降时才接受候选记忆,为随机评估提供置信度门控,并为分段平稳环境提供重新锚定保证。Stochastic Reflective Memory Ascent (SRMA), which accepts a candidate memory only after a grounded evaluation risk strictly decreases, is introduced and provides confidence gating for stochastic evaluation and re-anchoring guarantees for piecewise-stationary environments.

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction
onPanda:通过 token 级校正高效标注 LLM 与 Agent 的 on-policy 对齐数据。
arXiv:2609.24983 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 OnPanda,一种用于高效标注 LLM 对齐数据与 Agent 轨迹的交互式工具,以 token 级修正为核心交互方式;并发布使用 onPanda 标注的数据集 Panda-CVL,以及一个面向 token 级修正的基准。OnPanda is presented, an interactive tool for efficiently annotating LLM alignment data and agent trajectories that adopts token-level correction as its core interaction and releases Panda-CVL, a dataset annotated with onPanda, together with a benchmark for token-level correction.

Self-Organizing Agent Teams Learn to Reason Together
自组织 Agent 团队学习协同推理
arXiv:2609.22682 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文提出 Self-Organizing Agent Teams (SAT),即从先前协作中学习可复用策略、固定成员队伍的 AI 代理,用以组织角色、对话阶段、参与方式与信息流,表明组织本身可成为代理的一项能力。Self-Organizing Agent Teams (SAT) are introduced, fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow, suggesting that organization itself can become an agent capability.

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis
IterSynth:通过角色解耦的迭代合成重新思考深度搜索 Agent
arXiv:2609.29444 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 IterSynth,一种角色解耦、基于摘要的范式,在用于识别信息需求的 Planner 与用于将证据整合进演化摘要状态的 Synthesizer 之间交替,作为一种模型无关的提示范式,在前沿闭源模型上相对 ReAct 及类似提示范式取得显著的零样本增益。IterSynth is proposed, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information needs and a Synthesizer for integrating evidence into an evolving summary state, and serves as a model-agnostic prompting paradigm, delivering substantial zero-shot gains over ReAct and similar prompting paradigms on frontier proprietary models.

JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces
JevSpawn:通过组合动作空间实现自适应 Agent 推理
arXiv:2610.00437 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 JevSpawn,一种将自然语言任务规范连接到有限概率探索的组合策略,并将 JevSpawn 确立为结构化 Agent 推理的一种有前景的方法,在任务性能和导航速度上均有所提升。This work introduces JevSpawn, a compositional policy that connects natural language task specifications to finite probabilistic exploration, and establishes JevSpawn as a promising approach to structured agentic inference, with improved task performance and faster navigation.

Automating the Design of Embodied Agent Architectures
具身智能体架构设计的自动化
arXiv:2606.30111 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文在视觉语言导航、具身问答和语言条件操控任务上评估了三种 AAS 变体,覆盖四个具身执行器,结果表明架构级搜索能在具身任务上产生可部署且具有方向性的成功率提升,而其中一个看似得分较高的候选因存在泄漏而被判定为无效。This work evaluates three AAS variants across four embodied executors spanning vision-language navigation, embodied question answering, and language-conditioned manipulation, and shows that architecture-level search can produce deployable and directional success-rate gains on embodied tasks, while one apparent high-scoring candidate is rejected as leak-bearing.

ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition
ReDesign:通过 Agentic 分解从图像中恢复可编辑的设计结构
arXiv:2607.25565 Agent 智能体 观点 OA · 绿色 被引 2 · S2

本文提出 ReDesign,一种 agentic 框架,通过跨模态选择与组合专用工具来构建可编辑的层级结构(layer hierarchy),在取得强视觉保真度的同时,于布局、颜色与文本编辑上提供最高的可编辑性。ReDesign is presented, an agentic framework that grows an editable layer hierarchy by selecting and composing specialized tools across modalities, and achieves strong visual fidelity while delivering the highest editability across layout, color, and text edits.

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm
超越"感觉更好":能力维持型情感对话作为一种纵向研究范式
arXiv:2607.27851 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

情感对话研究包含两种颇具影响力的策略传统。共情对话优先理解说话者的情绪体验;情感支持对话则选择并排序以满足求助者当前的需求。持续使用引入了更进一步的目标:有效的支持应在整个交互生命周期中维持用户进行情绪调节、应对、自我认同决策以及社会联结的能力。我们提出能力维持型情感对话(CSED)作为一种纵向研究范式,将支持策略与上述目标对齐,并...Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience. Emotional support conversation selects and sequences support for the seeker's current needs. Sustained use introduces a further goal. Effective support should sustain users' capacities for emotion regulation, coping, self-endorsed decisions, and social connection across the interaction lifecycle. We propose capability-sustaining emotional dialogue (CSED) as a longitudinal research paradigm that aligns supportive strategy with this goal and or