AI Agent 工具调用 · 长期记忆 · 多代理协作 — 候选整理

整理时间: 2026-09-28(周一)
主题: AI Agent 工具调用、长期记忆、多代理协作轻量候选
数据来源: HF Daily 富化 + Substack 补充(arXiv 406 降级)


核心洞察(Substack 来源)

1. Agent 记忆体系的三层架构

  • 短时记忆:上下文窗口,单轮运行
  • 长期记忆:跨会话持久存储(客户偏好、累计上下文)
  • 共享记忆:多 Agent 之间信息传递(如 researcher → writer → editor 流水线)

来源:sidsaladi.substack.com — The Complete Guide to Building AI Agents in 2026

2. 多 Agent 协作已是 2026 默认范式

  • Gartner 预测 2027 年 1/3 的 agentic 部署将采用多 Agent
  • CrewAI、AutoGen 原生支持多 Agent;LangGraph 通过图节点协调
  • 核心挑战:同步、过期 embedding、所有权边界、版本冲突、检索污染

来源:futureagi.substack.com — Top 5 Agentic AI Frameworks to Watch in 2026

3. 多 Agent 共享记忆的深层陷阱

一个真实故障案例:多个 Agent 将合成摘要回写长期记忆后,生成性解释逐渐取代了源真值,系统开始从自身抽象中学习,输出仍保持连贯但已偏离事实。

来源:rockybhatia.substack.com — How to Learn Agentic AI in 2026

4. 2026 Agent 技术栈演化路线图

  • 2020–2022:无状态 LLM wrapper
  • 2023:工具调用 + 记忆 wrapper(LangChain、AutoGen)
  • 2024:图编排(LangGraph、CrewAI)
  • 2025:MCP 标准化工具访问 + 结构化记忆
  • 2026:协调式 Agent 集群(planner、memory、verifier 角色分工)

来源:aiagentssimplified.substack.com — The 2026 Path to Learning AI Agents

5. Agent Memory 关键产品一览

产品 定位 事件
Mem0 Agent 记忆层 获 A 轮 2400 万美元
Letta (MemGPT) 长期记忆 LLM 系统 获 1000 万美元种子
Zep Agent 记忆 SOTA 替代 Mem0 基准
LongMemEval 长期交互记忆 Benchmark 基准评测

来源:codingwithroby.substack.com — The 2026 AI Agent Stack, Drawn from Scratch


论文候选(按相关性排序)

★ 高价值

1. Coding Agents for Generalized Task and Motion Planning Problems
https://arxiv.org/abs/2609.30233 | HF votes: 11
摘要: 编码 Agent 能自动化 TAMP(任务与运动规划)过程,综合跨实例泛化程序。离散决策与几何/运动约束紧耦合是核心难点,研究 Agent 如何在此类问题上实现自动化工程。标签:agent


2. Your Transformer Can Hold Two Thoughts at Once: Linear Superposition in LLMs
https://arxiv.org/abs/2609.29845 | HF votes: 75
摘要: Transformer 对线性组合输入产生叠加输出分布(Superposition Linearity Hypothesis)。该特性随预训练递减,但可通过轻量微调恢复。对理解 Agent 记忆编码机制有潜在意义。标签:research


○ 补充候选

3. Learning to Discover Interesting Mathematics
https://arxiv.org/abs/2609.28603 | HF votes: 11
Agent 数学推理与有趣性判断能力相关,非直接工具/记忆方向。

4. Parts-of-Speech as Emergent Categories in SAE Latent Space
https://arxiv.org/abs/2609.29362 | HF votes: 10
SAE 可解释性研究,与 Agent 内省工具有弱关联。

5. Just Ask Jev: RLCD for Zero-Shot Alignment Failure Detection
https://arxiv.org/abs/2609.29429 | HF votes: 7
对齐失败检测 Benchmark(RLCDAlignBench),Agent 安全评测方向。

6. RGBD20K
https://arxiv.org/abs/2609.29028 | HF votes: 8
RAG benchmark 相关性低,纯视觉任务。

7. AV-GRPO: Joint Audio-Video Diffusion RL
https://arxiv.org/abs/2609.29816 | HF votes: 5
多模态生成强化学习,Agent 工具集扩展方向。

8. DeltaWAM: World Action Models for Bimanual Manipulation
https://arxiv.org/abs/2609.28811 | HF votes: 4
世界-动作模型,机器人控制方向,工具使用泛化相关度低。


总结

本次候选共 8 条,高价值 2 条(均来自 HF Daily),Substack 补充洞察 5 条,CSDN 未使用。

arXiv 搜索因 406 降级,论文候选质量受限。Substack 内容提供了 2026 Agent 技术栈、记忆产品图谱和多 Agent 共享记忆风险的真实案例,弥补了论文侧的不足。

下周建议: 优先修复 arXiv API query,增加 "MCP"、"agent memory benchmark"、"multi-agent coordination evaluation" 专项检索。