Papers · organized/paper_cards

论文

267 张论文卡片 · Agent 智能体

开放获取 全部 绿色 · 724
RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents
RODS:面向多轮工具使用 Agent 的奖励驱动在线数据合成
arXiv:2606.19047 Agent 智能体 方法 OA · 绿色 被引 1 · S2

RODS(Reward-driven Online Data Synthesis)通过将进度奖励方差重新用作零成本边界检测器,在 RL 训练与数据生成之间形成闭环,无需在训练已有的 rollout 之外增加额外推理。RODS (Reward-driven Online Data Synthesis) closes the loop between RL training and data generation by repurposing the progress reward variance as a practical, zero-cost boundary detector that requires no extra inference beyond the rollouts already computed for training.

TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction
TRAP:任务完成与对主动隐私窃取抵抗力的基准
arXiv:2606.18996 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

对涵盖前沿闭源与开源模型、共 22 个模型在多个规模上的评估发现,所有模型家族均存在不可忽略的隐私泄露,且指令遵循能力与泄露率呈正相关。Evaluating 22 models spanning frontier proprietary and open-source models at multiple scales, it is found that all model families exhibit non-trivial leakage, and that instruction- following ability correlates with leakage rate.

Searching for Synergy in Shared Workspace Human-AI Collaboration
在共享工作空间的人机协作中寻找协同效应
arXiv:2606.18413 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作以模拟的共享工作空间人机团队为受控实验环境,研究协作结构如何影响团队行为,并表明协调结构是决定可用能力能否提升团队结果的关键。This work uses simulated shared-workspace human-AI teams as a controlled testbed for studying how collaboration structure shapes team behavior, and suggests that coordination structure is central to whether available capability improves team outcomes.

ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
ProvenanceGuard:面向 MCP-Based LLM Agent 的来源感知事实性验证
arXiv:2606.18037 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,在基于 MCP 的 Agent 中,来源归因是事实性验证的一个独立维度;提出 ProvenanceGuard,一种针对 MCP 依据回答的来源感知验证器。Results show that source attribution is an independent axis for factuality verification in MCP-based agents, and ProvenanceGuard, a source-aware verifier for MCP-grounded answers is introduced.

Cordon: Semantic Transactions for Tool-Using LLM Agents
Cordon:面向工具调用 LLM Agent 的语义事务
arXiv:2606.17573 Agent 智能体 方法 OA · 绿色 被引 5 · S2

本文介绍 Cordon,一个事务性运行时系统,用于在提交前暂存并验证 Agent 的不可逆操作;其在保持良性任务完成的同时降低不可逆操作失败率,且仅带来适度的审批与时延开销。This paper introduces Cordon, a transactional runtime system for staging and validating irreversible agent effects before commit and reduces irreversible-effect failures while preserving benign task completion with modest approval and latency overhead.

Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
你的 AI 旅行 Agent 会为你预订一场斗牛:面向前沿 AI 模型的隐式动物福利 Agent 基准
arXiv:2606.18142 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

结果表明模型倾向于选择有害场景,在中性预订选项上的表现低于随机猜测水平,其中 Claude 4.8 取得最高分 64.7%。The results indicate that models tend to prefer harmful scenarios, performing below the random chance rate for selecting a neutral booking option, with Claude $4.8$ achieving the highest performance at $64.7\%$.

Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications
基于 Agentic AI 的医疗应用中过早诊断交接与静默幻觉缓解框架
arXiv:2606.18068 Agent 智能体 应用落地 OA · 绿色 被引 1 · S2

提出一种多 Agent 框架,通过以确定性编排约束替代 "LLM-as-a-judge" 路由,解决可能在到达患者前未被发现的过早诊断交接与静默临床幻觉问题;观察到 OLDCARTS 完整度与语义熵之间存在统计显著的负相关,提示结构化信息采集与诊断不确定性降低相关。A multi-agent framework that addresses premature diagnostic handoff and silent clinical hallucinations that may go undetected before reaching the patient by replacing ``LLM-as-a-judge''routing with deterministic orchestration constraints is proposed and observes a statistically significant negative correlation between OLDCARTS completeness and semantic entropy, suggesting that structured information gathering is associated with reduced diagnostic uncertainty.

Environment-Grounded Automated Prompt Optimization for LLM Game Agents
面向 LLM 游戏 Agent 的环境接地自动化提示优化
arXiv:2606.17838 Agent 智能体 方法 OA · 绿色 被引 2 · S2

提出一种针对 LLM Agent 的自动化提示优化框架,将"观测到动作"流水线分解为目标条件描述子 Agent 与动作选择 Agent,并通过由 LLM 驱动、以环境回报为指导的进化循环迭代优化各模块提示。An automated prompt optimization framework for LLM agents that decomposes the observation-to-action pipeline into a goal-conditioned descriptor agent and an action selection agent, and iteratively refines each module's prompt through an LLM-driven evolutionary loop guided by environment returns is introduced.

The Integrator Advantage: Controlled Agentic AI for Small and Medium-Sized Companies
集成商优势:面向中小型企业的受控 Agentic AI
arXiv:2606.16649 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文认为,Agentic AI 的近期价值不在于完全自主或削减人力,而在于面向简单与中等复杂度业务流程的受控部分自主。It is argued that the near term value of Agentic AI does not lie in full autonomy or workforce reduction, but in controlled partial autonomy for simple and medium complexity business processes.

When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting
当 Agent 自动化变得有利可图:通过 Trace-Economic Underwriting 量化并承保自主 AI 风险
arXiv:2606.16465 Agent 智能体 应用落地 OA · 绿色 被引 8 · S2

本文引入 trace-economic underwriting,将工具调用 trace 映射为客户风险敞口与可索赔损失,并以此表示用于定价、控制与风险转移,使用确定性经济标签而非 LLM 评判器。T trace-economic underwriting is introduced, which maps tool-use traces to customer exposure and claimable loss, then uses this representation for pricing, control, and risk transfer, and uses deterministic economic labels rather than an LLM judge.

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
现实场景下工具调用型 LLM Agent 数据泄露风险评估
arXiv:2606.17114 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

新加坡 AI Safety Institute 与韩国 AI Safety Institute 联合评估了涵盖客服、DevOps、网页自动化以及企业与个人生产力场景下 12 项真实非对抗任务中的 Agent 数据泄露问题,表明操作性数据泄露是与对抗性数据外泄不同的一阶 Agent 安全问题。A joint evaluation by the Singapore AI Safety Institute and the Korea AI Safety Institute examining agent data leakage in 12 realistic, non-adversarial tasks spanning customer support, DevOps, web automation, and enterprise and personal productivity indicates that operational data leakage is a first-order agent-safety concern distinct from adversarial exfiltration.

CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
CoffeeBench:面向异构多 Agent 经济中的长视野 LLM Agent 基准测试
arXiv:2606.16613 Agent 智能体 评测集 OA · 绿色 被引 6 · S2

对 Agent 行为的分析揭示了长视野经济交互中的显著差异:表现更好的模型与其他企业的沟通更为活跃,而 Claude Haiku 4.5 则表现出 idle-drift 失效模式,在生成连贯评估与规划的同时仍反复选择不行动。Analysis of agent behavior reveals substantial differences in long-horizon economic interaction: higher-performing models communicate more actively with other firms, whereas Claude~Haiku~4.5 exhibits an idle-drift failure mode, repeatedly choosing inaction despite producing coherent assessments and plans.

AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents
AgentOdyssey:面向测试时持续学习智能体的开放式长视野文本游戏生成
arXiv:2606.24893 Agent 智能体 方法 OA · 绿色 被引 1 · S2

AgentOdyssey 被提出,这是一个新颖的评估框架,通过程序化方式生成包含丰富实体、世界动态和长视野任务的开放式文本游戏,并发现短期记忆对多种智能体范式均有益,是智能体测试时训练的重要组成部分。AgentOdyssey is introduced, a novel evaluation framework that procedurally generates open-ended text games with rich entities, world dynamics, and long-horizon tasks and finds that short-term memory benefits multiple agent paradigms and is an important component of agent test-time training.

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents
超越攻击成功率:面向使用工具的 AI 智能体的动作分级严重性量表
arXiv:2607.07474 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出了一种动作分级伤害评估量表,按七级有序量表对智能体的工具调用轨迹进行打分,评判依据包括所执行动作的可逆性、是否越界涉及其他方以及是否扩大了权限。An action-graded harm rubric is introduced that scores an agent's tool-call trajectory on a seven-level ordinal scale according to whether the executed action was reversible, whether it crossed scope to reach another party, and whether it expanded privilege.

Automating the Design of Embodied Agent Architectures
具身智能体架构设计的自动化
arXiv:2606.30111 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文在视觉语言导航、具身问答和语言条件操控任务上评估了三种 AAS 变体,覆盖四个具身执行器,结果表明架构级搜索能在具身任务上产生可部署且具有方向性的成功率提升,而其中一个看似得分较高的候选因存在泄漏而被判定为无效。This work evaluates three AAS variants across four embodied executors spanning vision-language navigation, embodied question answering, and language-conditioned manipulation, and shows that architecture-level search can produce deployable and directional success-rate gains on embodied tasks, while one apparent high-scoring candidate is rejected as leak-bearing.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
面向 Agentic 强化学习的单 Rollout 异步优化
arXiv:2607.07508 Agent 智能体 方法 OA · 绿色 被引 6 · S2

提出 Single-rollout Asynchronous Optimization(SAO),用于解决异步 RL 中的稳定性与 off-policy 难题,可稳定训练一千步,并在 Agentic 编码与推理基准上一致优于 GRPO 及其变体。Single-rollout Asynchronous Optimization (SAO) is presented to address the stability and off-policy challenges in asynchronous RL and is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks.

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
AgentLens:面向编码 Agent 评估的生产级轨迹评审
arXiv:2607.06624 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

提出 AgentLens,一个面向交互式代码 Agent 的生产级评估基准,将形式化验证(在存在客观检查时)与 LLM 编写的轨迹评审及并排比较相结合,使每次运行都能给出关于分数为何如此的可读解释。This work presents AgentLens, a production-assessed benchmark for interactive code agents that pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is.

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
UniClawBench:面向真实世界任务的主动 agent 通用基准
arXiv:2607.08768 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

提出首个面向动态真实世界场景评估主动 agent 的能力驱动基准,并在模型与框架层面的全面比较表明,基模型能力与 agent 框架设计共同决定了真实世界环境中的性能表现。The first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world settings is introduced, and comprehensive comparisons across both models and frameworks show how base model capabilities and agent framework designs jointly shape performance in real-world environments.

Token-Flow Firewall: Semantic Runtime Auditing for Persistent AI Agents
Token-Flow Firewall:面向持久化 AI Agents 的语义运行时审计
arXiv:2607.08395 Agent 智能体 方法 OA · 绿色 被引 1 · S2

提出 TokenWall,一种作用于 agent token 流的语义防火墙式运行时防御框架,证明语义运行时约束可在持久化 AI agents 上实现实用的安全性与效用性权衡。TokenWall is proposed, a runtime defense framework that acts as a semantic firewall over agent token flows, demonstrating that semantic runtime containment can achieve a practical security-utility trade-off for persistent AI agents.

The Context Access Divide: Interaction-Level Architecture as a Complementary Dimension of Agentic Inequality
上下文访问鸿沟:交互级架构作为 Agent 不平等的一个补充维度
arXiv:2607.08495 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Contextuality(上下文性)——即 AI 系统自主访问用户累积知识资本的程度——作为 AI 介导不平等的一个维度,补充但不可化约为 Sharp 等人的框架。Contextuality -- the degree to which an AI system autonomously accesses a user's accumulated knowledge capital -- is proposed as a dimension of AI-mediated inequality that complements, but is not reducible to, the Sharp et al. framework.

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
Remember When It Matters:面向长视野 Agent 的主动记忆 Agent
arXiv:2607.08716 Agent 智能体 方法 OA · 绿色 被引 1 · S2

消融实验表明,选择性干预优于被动记忆库暴露、常驻注入、仅顾问引导和通用检索。Ablations show that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, advisor-only guidance, and general retrieval, and general retrieval and that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, and general retrieval.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
LongMedBench:面向长程临床决策的医疗 Agent 基准测试
arXiv:2607.09322 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 LongMedBench,一个基于真实 EHR 的长程临床决策基准,并设计了一套包含三个评估维度的分类体系:事实型问答、时序推理、长程决策。This work introduces LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making, and proposes an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Long-Horizon-Terminal-Bench:基于密集奖励评分的 Agent 长程终端任务极限测试
arXiv:2607.08964 Agent 智能体 评测集 OA · 绿色 被引 7 · S2

该工作提出 Long-Horizon-Terminal-Bench,一个涵盖九大类共 46 个长程任务的终端基准,包括实验复现、软件工程、多模态分析、交互游戏与科学计算,并分析失败模式与错误规律,以推动长程终端 Agent 的后续研究。This work introduces Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing, and analyzes failure modes and error patterns to support future progress on long-horizon terminal agents.

Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation
Flow-ERD:面向多样化交通仿真的 Agent 类型感知 Flow Matching 与熵正则蒸馏
arXiv:2607.06957 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

论文提出 Flow-ERD,一个同时追求真实性与多样性的多 Agent 仿真器,在 WOSAC 测试基准上排名第一,并在可复现基线中主导真实性-多样性 Pareto 前沿。Flow-ERD is introduced, a multi-agent simulator that pursues realism and diversity jointly and ranks first on the WOSAC test benchmark and dominates the realism--diversity Pareto front among reproducible baselines.

Metacognition in LLMs: Foundations, Progress, and Opportunities
LLM 中的元认知:基础、进展与机遇
arXiv:2607.11881 Agent 智能体 综述 OA · 绿色 被引 1 · S2

本文首次全面综述了 LLM 元认知的研究现状,涵盖用于测量和评估 LLM 元认知能力的方法与 benchmark、激发、改进与应用 LLM 元认知的技术,以及当前研究的发现与启示。The first comprehensive overview of the current state of knowledge on metacognition for LLMs is presented, including methods and benchmarks to measure and evaluate LLMs'metacognitive abilities, techniques to elicit, improve, and apply metacognition in LLMs, and findings and implications of ongoing research.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
ABot-AgentOS:具备终身多模态记忆的通用机器人 Agent 操作系统
arXiv:2607.10350 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ABot-AgentOS,一个通用机器人 Agent Operating System,位于底层控制器之上,提供 deliberation agent 层,支持场景条件规划、上下文隔离的 Skill 执行、多阶段验证、多模态记忆以及边云协同。ABot-AgentOS is presented, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration.

Multi-Agent LLMs Fail to Explore Each Other
Multi-Agent LLMs 未能互相探索
arXiv:2607.11250 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 Multi-Agent Contextual Exploration (MACE),一个通过结构化的对等体选择显式促进探索的轻量级框架,显著改善了探索行为和下游任务表现,并在理论上证明探索价值随 Agent 多样性增加而提升。This work introduces Multi- Agent Contextual Exploration (MACE), a lightweight framework that explicitly promotes exploration through structured peer selection that substantially improves exploration behavior and downstream task performance and shows theoretically that the value of exploration increases with agent diversity.

Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution
修复前先知:面向软件问题解决的 QA 驱动仓库知识获取
arXiv:2607.11111 Agent 智能体 方法 OA · 绿色 被引 2 · S2

基于 LLM 的编程 Agent 显著推动了自动化软件问题解决,但由于对仓库理解不足,仍易出现事实性错误。近期方法尝试通过修复前仓库探索来缓解此问题;然而,其修复驱动策略在未识别 Agent 知识缺口的情况下探索仓库,往往产生不精确的上下文,无法弥补潜在的理解不足。本文提出 ACQUIRE,一种面向软件问题解决的 QA 驱动框架,模拟经验丰富的开发者LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods attempt to mitigate this limitation through pre-repair repository exploration; however, their fix-driven strategies explore repositories without identifying the agent's knowledge gaps, often yielding imprecise context that fails to bridge the underlying understanding deficit. In this paper, we propose ACQUIRE, a QA-driven framework for software issue resolution. Mirroring how experienced developer

Towards Autonomous and Auditable Medical Imaging Model Development
迈向自主且可审计的医学影像模型开发
arXiv:2607.10522 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 AMID,一种面向医学影像模型开发的自主多 Agent 框架,其性能优于所评估的通用 MLE 系统,并在异构任务上接近或匹配强大的人工设计挑战赛方案。AMID is introduced, an autonomous multi-agent framework for medical imaging model development that outperformed evaluated general-purpose MLE systems and approached or matched strong human-designed challenge solutions across heterogeneous tasks.

Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
面向 Coding Agent 基础模型的 Function-Aware Fill-in-the-Middle 中期训练
arXiv:2607.12463 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

除领域内增益外,mid-training 还能缓解 agentic post-training 对非 Agent 编程及非编程工具调用基准(tau-bench、BFCL)造成的能力侵蚀:尽管 mid-training 语料仅含 Python 代码,函数调用的归纳偏置在 post-training 后依然保留,带来稳定的增益。Beyond in-domain gains, mid-training mitigates the capability erosion that agentic post-training otherwise inflicts on non-agent coding and non-coding tool-use benchmarks (tau-bench, BFCL): although the mid-training corpus contains Python code only, the function-call inductive bias survives post-training and yields consistent gains.

Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
Navigating the Mirage:面向鲁棒误导性图表问答的双路径 Agentic 框架
arXiv:2603.28583 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

尽管视觉-语言模型(VLMs)已取得成功,但误导性图表因欺骗性视觉结构与失真数据表示仍构成重大挑战。我们提出 ChartCynics,一个通过"怀疑式"推理范式揭露视觉欺骗的 Agentic 双路径框架。与整体化模型不同,ChartCynics 将感知与验证解耦:诊断式视觉路径通过策略性 ROI 裁剪捕获结构异常(如倒置坐标轴),OCR 驱动数据路径确保数值根植性。为解决跨模态冲突,我们提出Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework designed to unmask visual deception via a "skeptical" reasoning paradigm. Unlike holistic models, ChartCynics decouples perception from verification: a Diagnostic Vision Path captures structural anomalies (e.g., inverted axes) through strategic ROI cropping, while an OCR-Driven Data Path ensures numerical grounding. To resolve cross-modal conflicts, we introduce

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
Search Beyond What Can Be Taught:Agentic 视觉生成中的知识边界演化
arXiv:2607.05382 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本研究将朴素搜索的根因追溯到生成器特有的、可演化的知识边界——即生成器经训练可内化的内容与必须保留于外部上下文的内容之间的鸿沟,并表明该边界可通过"先教后搜"协同训练框架被有效发现。This work traces the root cause of naive search to a generator-specific, evolving knowledge boundary: the divide between what a generator can internalize through training and what must remain in external context, and shows that it is discoverable through a teach-then-search co-training framework.

When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects
When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects
arXiv:2607.13679 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本研究考察了 GitHub 项目在各自引入首个 bot 前后两年间的情况,发现变化集中在采纳时点附近,而非逐渐累积,这与一种特定解读一致:可预测、基于规则的 Agent 能够成为社区社交基础设施的一部分。This work examines GitHub projects for two years before and after each adopted its first bot, finding changes cluster around adoption rather than accumulating gradually, consistent with a specific interpretation: predictable, rule-based agents can become part of a community's social infrastructure.

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
arXiv:2607.11523 Agent 智能体 方法 OA · 绿色 被引 2 · S2

论文提出 Vinci2,一个主动式的第一人称视频协助系统,将端侧助手 Vinci 由被动响应推进到主动协助;以及免训练、记忆增强的 Agent EgoMemo,维护三种互补的记忆表征:多尺度时间摘要、语义知识图谱与视觉嵌入档案。Vinci2 is presented, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity and EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives.

Tracing Agentic Failure from the Flow of Success
Tracing Agentic Failure from the Flow of Success
arXiv:2607.12747 Agent 智能体 方法 OA · 绿色 被引 2 · S2

论文提出 OAT,将该问题建模为基于神经受控微分方程的单类学习,在潜空间中刻画成功轨迹的动力学模式;实验表明其比基于 prompt 的基线更快,并在领域内和分布外数据集上均稳定优于基线。OAT is proposed, which casts this problem as one-class learning with neural controlled differential equations, modeling the dynamical pattern of successful trajectories in latent space, and is shown to be faster than prompting-based baselines and consistently outperforms them in both in-domain and out-of-distribution datasets.

From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization
From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization
arXiv:2607.07702 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

STRACE(Structural TRajectory Analysis and Causal Extraction)是一个用于构建高信噪比优化上下文的框架,旨在对长周期 Agent 实施更精确、更有效的优化。STRACE (Structural TRajectory Analysis and Causal Extraction) is a framework that constructs high signal-noise optimization contexts for more precise and effective optimization of long-horizon agents.