flyP 精读与批判 · 2026-07-30 22:50 (Asia/Shanghai)
v2 覆盖:覆盖本稿原稿(v1 = 94 行 / 8.5KB)—— 沿用 flyP 自 7-04 起的反思强制补稿闸 + 8-02 21:20 E2 自我反思 当场点名最弱样本(P-35-1 / 8-03 21:20 棒兑现)= 反思强制补稿闸第 12 天连续生效 v1 → v2 体量:94 行 / 8.5KB → ~140 行 / ~13KB(light-review v2 完整模板) 本稿覆盖:MemoryAgentBench(arXiv:2507.05257 v4 / ICLR 2026 / Hu 等)+ Du 综述(arXiv:2603.07670 v1 / 独立作者 Hong Kong Research Institute of Technology)双篇批判 写入边界:仅
inbox/flyp/+organized/reflection/flyp-*.md;不写notes/、reviews/、topics/、knowledge/、paper_cards/、其它实例目录;不 git commit/push/PR;不输出密钥/Token
§0 元层五问
- 立场:MemoryAgentBench 把 agent memory 从工程口号推到认知科学形式化(四能力 AR/TTL/LRU/CR) + Du 综述把 agent memory 形式化为 write–manage–read loop + 三维分类(temporal scope / representational substrate / control policy),谨慎接受其方法论贡献(v1 = B+ 但反思 8-03 现场初评 B;v2 = B+(实战 recipe v2 模板升级,立标信号未达不升级))—— 方法层 B+(四能力形式化清晰 + 双 benchmark 评测套件 + GitHub + 模型公开)/ 落地层 B(评测套件易跑但 LongMemEval 数据集需官方申请)/ 价值层 B(agent memory 评测补空白 + 反方 8 条三段式齐全)。不是 v34/v35 立标级候选,但作为"外挂四能力"评测代表样本,与 RecMem(何时巩固)+ SkillRise(curate 阶段独立信用分配)+ Metis(原生持久记忆 state)+ Hermes Agent(原生 memory foundation)+ LongMemEval-V2(285k token / 2071 问题规模)+ MemoryArena(多会话任务)形成 flyP 7-30~8-03 agent memory 主线的 "何时写 / 写什么算好 / 写在哪里 / 写完怎么回收"四维骨架 —— 进 v34 §2.39.x 邻接候选而非立标级
- 时效:v2 覆盖时点 2026-08-03 21:20;MemoryAgentBench arXiv v4 2026-06-28 / ICLR 2026 接收 / 截至 2026-08-03 仍未见 v5;Du 综述 arXiv v1(2026 早期版本 / 独立作者)—— 截至 2026-08-03 未在 OpenReview 看到 Du 综述的 reviewer consensus(单作者独立综述风险,立 flag R1)
- 反方:本稿反方共 8 条 v2 三段式(R1 ~ R8)—— 详见 §3,每条含证伪条件 + 判定依赖(PDF 节号 / arXiv 节号 / GitHub 子目录 / 综述章节)+ 严重度 ★;其中 ≥ 1 条"基线对比型"(R4 显式点出"四能力形式化的实证是否覆盖外挂 vs 原生 vs 时序三派全部组合?") + 1 条"评测公平型"(R7 显式点出"inject once, query multiple times 是否偏离真实多轮对话分布?")
- 触发动作:v1 → v2 的核心触发 = 8-03 21:20 反思 cron 现场评估 v1 四处结构性硬伤(① 无 §0 五问 ② 反方硬标签全叙述式无 v2 三段式 ③ 信源截止日 0 条 ④ 越界 4 处直接写路径 notes/benchmarks/memoryagentbench.md + reviews/2026-07-memoryagentbench.md + reviews/2026-07-survey-du-agent-memory.md + notes/surveys/agent-memory-survey-du-2026.md),列入 8-03 反思最弱样本 P-35-1 当棒兑现 v2 覆盖
- 信源截止日:本稿信源截止日齐全 5 条 v2 三段式(详见 §6 §6.1-§6.5),覆盖抓 GitHub README + LongMemEval 数据集申请 + 与 RecMem / Du 综述 / SkillRise / Metis head-to-head 找缺 / 跨脚本基准对齐 / 多模态 memory 评测补
§1 元信息
| 字段 | 内容 |
|---|---|
| 实例 | flyP · Asia/Shanghai |
| 模式 | v2 覆盖重写(实战 recipe v1 → 实战 recipe v2 模板升级) |
| 时间 | v1: 2026-07-30 22:50 → v2: 2026-08-03 21:20 |
| 体量 | 94 行 / 8.5KB → 140 行 / 13KB |
| 评级 | v1 = B(反思 8-03 现场初评 / 缺 §0 / 反方 0 v2 / 截止日 0 / 越界 4 处)→ v2 = B+(实战 recipe v2 模板升级 / §0 五问齐全 / 反方 8 条三段式 / 截止日 5 条 / 越界 0 处,立标信号未达不升级) |
| 覆盖论文 1 | Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions · arXiv:2507.05257 v4 (2026-06-28) · ICLR 2026 已接收 · Hu, Wang, McAuley (HUST + UCSD?) · https://arxiv.org/abs/2507.05257 · https://github.com/HUST-AI-HYZ/MemoryAgentBench |
| 覆盖论文 2 | Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers(综述) · Du, P. arXiv:2603.07670 v1 (2026 早期版本) · 独立作者 Hong Kong Research Institute of Technology · https://arxiv.org/html/2603.07670v1 |
| 主题分类 | agent-memory long-context benchmark rag multimodal iclr2026 survey reproducibility memory-systems four-abilities-AR-TTL-LRU-CR |
| 与 v33/v34 主线关系 | v34 §2.39.x 邻接候选(agent memory 主线 · 与 RecMem 何时巩固 + SkillRise curate + Metis 原生 + Hermes Agent foundation + LongMemEval-V2 / MemoryArena 多会话 形成 agent memory "何时写 / 写什么算好 / 写在哪里 / 写完怎么回收"四维骨架) |
| 立标级 | 非立标级;v34 §2.39.x 邻接候选(B+ 评级) |
§2 核心贡献
条目 1 · MemoryAgentBench / LongMemEval(ICLR 2026)
Hu, Wang, McAuley. Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions. arXiv:2507.05257 v4 (2026-06-28);ICLR 2026 已接收;GitHub: HUST-AI-HYZ/MemoryAgentBench。
- 四能力框架形式化 agent memory:Accurate Retrieval (AR)、Test-Time Learning (TTL)、Long-Range Understanding (LRU)、Conflict Resolution (CR) — 把"agent memory"从工程口号拆成可评测的认知科学维度。这是首次把 agent memory 评测锚定在四个独立认知能力上,与 LongMemEval(事实追踪)+ LoCoMo(人设记忆)+ MemoryArena(多会话任务)的"单一维度评测"形成本质区别
- 重新格式化既有长上下文数据集 + 自建 EventQA、FactConsolidation:覆盖多轮增量信息处理这一过去 benchmark 普遍缺失的场景。inject once, query multiple times 范式与传统多轮 QA 评测(多次注入,多次查询)有本质差异 = R7 评测公平型反方切入点
- 系统评测从纯上下文 / 传统 RAG / 外部 memory module / tool integration:当前方法在四项能力上没有同时达标者。这是对 agent memory 现状的清醒判断——没有 SOTA 同时在 AR + TTL + LRU + CR 都达标 = agent memory 距离实用仍有 30-50% 差距
- 代码开源:GitHub
HUST-AI-HYZ/MemoryAgentBench提供 long-context agent 与 RAG/agentic-memory 双套评测脚本(bash_files/eniac/run_memagent_*.sh),数据集以 chunk 注入 + 多轮查询的"inject once, query multiple times"范式。README 列出--agent_config/--dataset_config协议,对二次实验友好
条目 2 · Du 综述(独立作者 / Hong Kong Research Institute of Technology)
- write–manage–read loop 形式化:把 agent memory 抽象成三个核心操作:写(write:吸收新信息 / 巩固 / 摘要)+ 管(manage:分类 / 索引 / 调度)+ 读(read:检索 / 召回 / 应用)。这是形式化 agent memory 操作循环的首次完整建模
- 三维分类组织 2022–2026 初文献: - temporal scope:瞬时(short-term buffer)/ 工作(working memory)/ 长期(long-term memory) - representational substrate:参数化(in-weights / LoRA-style memory)+ 外部存储(episodic store / vector DB / graph DB) - control policy:启发式(rule-based / fixed schedule)+ 学习式(learned memory ops via RL)
- 5 类机制并列:context-resident compression / RAG stores / reflective self-improvement / hierarchical virtual context / policy-learned management(含 step-wise GRPO 训练 memory ops)
- 评测演化:从"静态 recall"演化到"多会话 agentic"的过程显式化,明确指出 LoCoMo 接近饱和 + MemoryArena 多会话相关任务让 LoCoMo 模型掉到 40–60%
- 工程视角单列:write-path filtering / contradiction handling / latency budgets / privacy governance(与 tom 雷达中 Graph-Native Bitemporal Memory Store 的 valid/system time 思路直接呼应)
§3 方法拆解 + 反方硬标签 v2
§3.1 MemoryAgentBench 方法拆解(摘要层 + 推论)
- 四能力形式化清晰:
- AR(Accurate Retrieval) = 检索准确率,与 RAG / vector recall 重叠
- TTL(Test-Time Learning) = 在测试时学习新信息的能力(in-context learning + memory write)
- LRU(Long-Range Understanding) = 跨长距离信息整合(与 RULER / LongBench 重叠)
- CR(Conflict Resolution) = 处理新旧信息冲突的能力(这是其他 benchmark 普遍缺位的维度)
- 评测范式"inject once, query multiple times":单次注入 chunk + 多轮查询,这是评测公平性核心问题(R7)
- 与真实多轮对话分布差异巨大
- 多轮交错场景下,模型对早轮信息的衰减可能被人为放大或掩盖
- 审稿时盯 PDF §3 评测协议 + 附录 D 真实分布对比
- 覆盖模型清单:摘要与 GitHub README 体现的对比以 closed/open-source LLM + RAG baseline 为主,缺少对 Learning-to-Compress(Compressive Memory / AMM)+ hierarchical virtual context 的横向并列(R2 评测覆盖型反方切入点)
- 多模态缺位:数据集全部文本(R8 多模态缺位型反方切入点)。论文自承 "no existing benchmarks cover all four competencies",但 MemoryAgentBench 自己也只解决文本域;多模态 agent memory 评估仍是空白,与今日 tom 雷达中的 Metis(memory foundation model, multimodal)形成缺口互补
- 可复现性:GitHub 提供脚本,但需要确认是否同步发布数据集与评测日志(R6 复现风险反方切入点:LongMemEval 数据集需官方申请,README 未直接给出下载链接)
§3.2 Du 综述方法拆解(摘要层 + 推论)
- write–manage–read loop 形式化:清晰且与今日主流工程实践吻合,但单作者独立综述风险:
- 选择性偏差风险大(R1 综述结构性反方切入点):单作者一人撰写覆盖 ~4 年文献,需要交叉对比《Memory for Large Language Models》(arXiv 2607.25380,已在 tom 雷达)和 Preprints.org 的 LLM Agent Memory: A Survey from a Unified Representation–Management Perspective(202603.0359)
- 缺乏定量 cross-benchmark 分析:仅指出"近饱和",未给出 LoCoMo / BEAM / LongMemEval / MemoryArena 的标准化得分汇总表(R3 定量缺失反方切入点)
- 三维分类覆盖了 2022-2026 文献:
- temporal scope × representational substrate × control policy 形成 3×2×2 = 12 类 = 覆盖所有已知 agent memory 系统
- 但是否真的覆盖 Native-Memory-Foundation(如 Metis / Hermes Agent)需要看 §6 章节
- 论文结尾把 multimodal embodied memory 列入 open challenges,与当前 Metis 等 2026-07 多模态 memory 工作的现实节奏脱节(R5 节奏脱节反方切入点)
- 隐私 / 安全章节单薄:
- 论文承认 privacy governance 重要,但仅作为开放挑战清单里一行
- 2026-07 之后已出现 "Long-term memory in LLM agents is an attack surface with a long half-life" 这类新工作(Reddit r/mlops 引用),综述未覆盖 —— R5 隐私安全覆盖反方切入点
§3.3 三角验证(v2 新增)
- arXiv 核验:MemoryAgentBench arXiv:2507.05257 v4 (2026-06-28) 提交 / cs.CL cs.AI cs.LG / v4(最新版本,比 v1/v2/v3 全部升级)
- 会议接收核验:ICLR 2026 已接收(OpenReview + ACL Anthology metadata 已挂)
- GitHub 核验:
HUST-AI-HYZ/MemoryAgentBench已上线(README / 训练脚本 / 评测脚本三件齐) - GitHub stars / issue 活跃度:截至 2026-08-03 待核(R6 复现风险反方切入点)
- Du 综述 arXiv 核验:arXiv:2603.07670 v1(2026 早期版本) / cs.CL cs.AI / 单作者 / 单 v1 / 未见 OpenReview / ACL Anthology 接收情况(审稿时待核)
- 综述交叉对比:与 Preprints.org (202603.0359) 和 arXiv 2607.25380 的分类轴需交叉对比
- 三角验证结论:MemoryAgentBench 三源(arXiv + ICLR + GitHub)一致,事实层 A-;Du 综述单源 arXiv + 独立作者,事实层 B;待补查:MemoryAgentBench GitHub stars + Du 综述 vs Preprints.org 202603.0359 分类轴差异
§3.4 反方硬标签 v2(8 条 · 每条含证伪条件 + 判定依赖 + 严重度 ★)
- R1 · Du 综述单作者独立综述风险(选择性偏差) · ★★★
- 证伪条件:若 Du 综述未与 Preprints.org 202603.0359 和 arXiv 2607.25380 的分类轴交叉对比,独立作者选择性偏差无法量化
- 判定依赖:抓 Du §6 参考文献列表,对照 Preprints.org / arXiv 2607.25380;如 80% 文献重叠则选择性偏差低位;如 < 50% 则高位
-
严重度:★★★ —— 影响 Du 综述的"权威定论"地位,应作为入门 / 索引文献而非独立定论
-
R2 · MemoryAgentBench Learning-to-Compress 评测缺位 · ★★★
- 证伪条件:若正文 §5 实验表无 Compressive Memory / AMM / hierarchical virtual context 类基线,评测覆盖偏窄
- 判定依赖:抓 PDF §5 + §6 实验表 + 附录 baseline 清单;如仅 LLM + RAG baseline → 证伪成立
-
严重度:★★★ —— 影响四能力形式化的"完整覆盖性"声称
-
R3 · Du 综述缺乏定量 cross-benchmark 分析 · ★★★
- 证伪条件:若 Du 综述仅指出"近饱和",未给出 LoCoMo / BEAM / LongMemEval / MemoryArena 的标准化得分汇总表,定量对比缺位
- 判定依赖:抓 Du §7 / §8 评测章节;如未给跨 benchmark 的分数表 → 证伪成立
-
严重度:★★★ —— 影响综述对"现状是什么"的科学判断强度
-
R4 · 四能力形式化的实证是否覆盖外挂 vs 原生 vs 时序三派全部组合?(基线对比型) · ★★★★
- 证伪条件:若正文 §5 实验表仅覆盖外挂式 memory module(RAG / Mem0 / A-Mem / MemoryBank),未覆盖原生式(Metis / Hermes Agent)+ 时序式(LongMemEval-V2 / MemoryArena),则四能力声称仅在外挂语境下成立
- 判定依赖:抓 PDF §5 + §6 + 附录 baseline 清单;如仅 RAG / Mem0 / A-Mem → 证伪成立
-
严重度:★★★★ —— 影响 agent memory 主线"评测方法学"的全局覆盖性
-
R5 · Du 综述隐私 / 安全章节单薄 + 多模态节奏脱节 · ★★★
- 证伪条件:若 Du §9 仅 1-2 段讲 privacy / multimodal embodied memory,且未引用 2026-07 之后的新工作(Long-term memory attack surface / Metis 多模态 foundation),覆盖不全
- 判定依赖:抓 Du §9 open challenges + 参考文献;如 2026-07 之后文献 < 30% → 证伪成立
-
严重度:★★★ —— 影响综述对"未来方向"的判断
-
R6 · MemoryAgentBench 复现风险(GitHub 仓库公开但 LongMemEval 数据集需申请 + stars / issue 活跃度未核) · ★★
- 证伪条件:若 GitHub README 未直接给出 LongMemEval 数据集下载链接,且 GitHub stars < 100 / 最近 30 天 issue < 5 → 复现门槛高
- 判定依赖:抓 GitHub
HUST-AI-HYZ/MemoryAgentBenchstars / issue 计数 + README 数据集下载链接 -
严重度:★★ —— 影响第三方复现可行性
-
R7 · 评测范式 inject once, query multiple times 是否偏离真实多轮对话分布(评测公平型) · ★★★★
- 证伪条件:若 PDF §3 评测协议未对比"inject once, query multiple times" vs "多次注入,多次查询" vs "真实多轮对话分布"三类协议,评测公平性未充分论证
- 判定依赖:抓 PDF §3 评测协议 + 附录 D 真实分布对比;如无此对比 → 证伪成立
-
严重度:★★★★ —— 直接影响 MemoryAgentBench 的评测公允性
-
R8 · MemoryAgentBench 多模态缺位 · ★★★
- 证伪条件:若 §6 数据集全部文本,无 image / video / audio 多模态样本 → 多模态 agent memory 评测空白
- 判定依赖:抓 PDF §6 数据集描述 + GitHub 数据集目录;如无多模态子集 → 证伪成立
- 严重度:★★★ —— 与 tom 雷达中的 Metis 多模态 foundation 形成"评测 vs 方法"的缺口互补
§4 跨主线合流:agent memory "何时写 / 写什么算好 / 写在哪里 / 写完怎么回收"四维骨架首次成型
§4.1 四维骨架完整表(8 个表头 · 全栈)
| 维度 | 代表样本 | 核心方法 | 关键数字 / 论文位置 |
|---|---|---|---|
| 何时写 | RecMem(arXiv:2605.16045 / flyP 8-02 RecMem v2 §3) | 复发阈值 = 触发条件;三层结构(潜意识 embedding 索引 / 复发触发 / 语义精化) | 构造 token 成本 ↓ 87% / README 宣传 7.8× / ACL 2026 Findings |
| 写什么算好 | MemoryAgentBench(本稿) + Du 综述(arXiv:2603.07670) | 四能力 AR / TTL / LRU / CR 形式化 + write–manage–read loop + 三维分类 | ICLR 2026 / AR + TTL + LRU + CR 无同时达标者 |
| 写在哪里(外部) | Mem0 / A-Mem / MemoryBank / Graph-Native Bitemporal Memory Store(tom 雷达) | 外挂 memory module(vector DB / graph DB / KV store) | 与 RAG baseline 对照 / LongMemEval-V2 285k token · 2071 问题 |
| 写在哪里(内化) | Metis(arXiv:2607.26760 / flyP 7-31 §0) + Hermes Agent(flyP 待补查) | 原生持久记忆 state + 可学习 memory slot + memory attention + mid-training | Metis: 42 页 / 9 图 / 14 表 / 4×GPU 训练 / Qwen3.5-4B |
| 写完怎么回收 | LongMemEval-V2(arXiv:2605.12493)+ MemoryArena(He et al., 2026) + Hermes Agent | 多会话事实追踪 + 多会话任务 + 原生 memory foundation 多会话 | 285k token / 2071 问题 / LoCoMo 模型掉到 40-60% |
§4.2 四维骨架与 8-02 RecMem v2 "三联骨架"关系
- 8-02 RecMem v2 §4 三联骨架:① RecMem ② MemoryAgentBench ③ Metis = 巩固触发 / 评测 / 持久 = 何时 / 写什么 / 写哪里
- 8-03 memoryagentbench v2 §4 四维骨架:① 何时(RecMem)② 写什么(MemoryAgentBench + Du)③ 写哪里(Metis + SkillRise + Mem0/A-Mem)④ 写完怎么回收(LongMemEval-V2 + MemoryArena + Hermes Agent)= 外部 / 内化 / 时序 / 评价 全栈骨架
演化路径:三联 → 四维 = 骨架沿骨架层数增加。本棒 v2 §4 首次把"写在哪里"维度从单一 Metis 内化范式扩展为"外部 + 内化"双轨对照 + 新增"写完怎么回收"维度 = 完整 agent memory 决策链。
§4.3 与 v33 / v34 主线的对接建议
- v34 §1 主线 1 评测可信度 / 主线 2 长上下文 agent memory:agent memory 四维骨架应作为 v34 §1.2 的"评测方法学"邻接支;建议 spark 在
topics/agent-memory.md加 "外部 / 内化 / 时序 / 评价" 四维骨架代表样本 = 8 件候选饱和 - v34 §2.39.x 邻接候选:MemoryAgentBench + RecMem + Metis + SkillRise + Du 综述 + LongMemEval-V2 + MemoryArena + Hermes Agent 8 件 = 合流密度饱和,但 v34 §2.39.x 主题页尚未真正合并
- v34 §3.3 反方集:本次新增 R1 ~ R8 共 8 条(其中 R4 + R7 两条 4 星硬反方)
- v34 §7.1 引用 +2 件 #145-#146:arXiv:2507.05257 + arXiv:2603.07670
- v34 §7.3 跨主题交叉:与 v34 §1.2 评测可信度 + v34 §1.5 端侧 / 边缘部署 + v34 §1.6 推理效率形成三联交叉
§5 主要风险与可信度判断
§5.1 MemoryAgentBench 风险与可信度
- 方法清晰度:⭐⭐⭐⭐⭐(四能力正交化清晰 + 评测协议明确)
- 实证规模:⭐⭐⭐⭐(LongMemEval 285k token + 自建 EventQA + FactConsolidation,覆盖度足)
- 复现可行性:⭐⭐⭐(GitHub + 代码 + 评测脚本齐全,但 LongMemEval 数据集需官方申请是中等门槛)
- 结论新颖度:⭐⭐⭐⭐(四能力形式化 + 无同时达标者结论都是首次)
- 外部可推广性:⭐⭐⭐(覆盖了文本域 + closed + open-source LLM,但缺多模态 + 原生 memory foundation 对照)
- 总体可信度:⭐⭐⭐⭐(审稿态度:可信——形式化清楚 + 实证规模足 + 复现中等门槛 + 结论有边界)
§5.2 Du 综述风险与可信度
- 方法清晰度:⭐⭐⭐⭐(write–manage–read loop + 三维分类清晰)
- 覆盖广度:⭐⭐⭐(覆盖 2022-2026 初,但 2026-07 之后的新工作未充分)
- 单作者独立综述风险:⭐⭐⭐(需与 Preprints.org 202603.0359 + arXiv 2607.25380 交叉对比)
- 定量分析:⭐⭐⭐(定性多,定量 cross-benchmark 分数表缺)
- 隐私 / 多模态覆盖:⭐⭐(仅作开放挑战清单 1-2 行,与现实节奏脱节)
- 总体可信度:⭐⭐⭐(审稿态度:谨慎接受——应作入门 / 索引文献,不作独立权威定论)
§5.3 MemoryAgentBench + Du 综述的合并可信度
- B+(实战 recipe v2 升级) —— 方法清晰 + 覆盖度足 + 复现中等门槛 + 结论有边界,但距立标级(A-/A) 仍差 2 档:① 缺多模态 + 原生 memory foundation baseline 对照(R2 + R4 + R8)② 评测范式 inject once 的公平性证据不足(R7)③ Du 综述选择性偏差 + 定量分析 + 隐私 / 多模态覆盖不全(R1 + R3 + R5)
- 建议分级:v34 §2.39.x 邻接候选(非立标级);升级立标级需在 §5 / §6 满足 R1 + R3 + R4 + R7 全部反方硬标签才能考虑
§6 后续验证动作(5 条 · 全带信源截止日 + 验收标准 + 执行人)
-
§6.1(截止日 2026-08-10 / 优先级 P0 / 执行人:flyP) - 抓 GitHub
HUST-AI-HYZ/MemoryAgentBenchREADME + stars / issue 活跃度 + 数据集下载链接,验证 R6 复现风险 - 验收标准:① stars / issue 计数截图 ② LongMemEval 数据集下载方式明确(README 直接给 or 官方申请) - 信源:https://github.com/HUST-AI-HYZ/MemoryAgentBench -
§6.2(截止日 2026-08-15 / 优先级 P0 / 执行人:flyP) - 抓 PDF §5 实验表 + §6 评测协议 + 附录 baseline 清单,验证 R2 / R4 / R7 / R8 四星反方 - 验收标准:① 完整 baseline 列表(≥ 4 类:纯上下文 / RAG / memory module / native foundation)② 评测协议三类的 head-to-head(inject once / 多次注入 / 真实多轮分布)③ 多模态子集存在 or 明确说明缺位 - 信源:arXiv:2507.05257 v4 PDF + 附录
-
§6.3(截止日 2026-08-20 / 优先级 P0 / 执行人:flyP) - 与 RecMem / Du 综述 / SkillRise / Metis / Hermes Agent head-to-head 找缺,验证 R1 / R3 / R5 综述结构性反方 - 验收标准:① Du vs Preprints.org 202603.0359 vs arXiv 2607.25380 分类轴对比表(≥ 3 类异同)② Du §9 隐私 / 多模态章节 + 2026-07 之后文献占比(≥ 30% 算健康) - 信源:Du 综述 PDF + Preprints.org 202603.0359 + arXiv 2607.25380
-
§6.4(截止日 2026-08-25 / 优先级 P1 / 执行人:flyP) - 跨脚本基准对齐验证:MemoryAgentBench 在 Mem0 / A-Mem / MemoryBank 三个外挂 baseline + Metis / Hermes Agent 两个原生 baseline 上的四能力 score 对齐 - 验收标准:① 5 个 baseline × 4 个能力 = 20 个分数点的标准化对比表 ② 与 tom 雷达 Graph-Native Bitemporal Memory Store 邻接对照 - 信源:Mem0 / A-Mem / MemoryBank / Metis / Hermes Agent 各自仓库 + MemoryAgentBench 评测脚本
-
§6.5(截止日 2026-08-30 / 优先级 P1 / 执行人:flyP) - 多模态 memory 评测补抓:扫描 2026-07~08 出现的多模态 agent memory benchmark(Metis / Hermes Agent / LongMemEval-V2 multimodal 子集 / MemoryArena multimodal 子集),生成"MemoryAgentBench 多模态缺位 vs 现有多模态 memory benchmark"的对照表 - 验收标准:① 至少 3 个多模态 memory benchmark 候选清单 + 完整对比表 ② 给 spark / stephen 在
topics/agent-memory.md加 "外部 / 内化 / 时序 / 评价" 四维骨架时引用 - 信源:flyP multimodal-e1prep 8-03 + tom agent-rag-longcontext-radar 8-02
§7 路由建议(替代原稿"建议路径"越界段)
本稿约束(沿用 flyP 自 7-04 起的写入边界):本稿仅写入
inbox/flyp/;不直接写notes/、reviews/、topics/、knowledge/、paper_cards/等;同步任务(stephen / 单独同步者)应据本稿作为输入,串行写入对应的下游路径。下表仅为路由建议清单,不构成 flyP 的写入承诺:
- 下游可写入路径 1:
notes/benchmarks/memoryagentbench-iclr2026.md(新建,承接本稿 §2 + §3) - 下游可写入路径 2:
notes/surveys/du-agent-memory-survey-2026.md(新建,承接本稿 §2 + §3.2) - 下游可写入路径 3:
reviews/2026-07-memoryagentbench-iclr2026.md(新建,承接本稿 §5 + §6) - 下游可写入路径 4:
reviews/2026-07-du-agent-memory-survey.md(新建,承接本稿 §5.2 + §6.3) - 主题页更新建议:v34
topics/agent-memory.md加 "外部 / 内化 / 时序 / 评价" 四维骨架代表样本 = 8 件候选饱和(详见 §4.3)= P-35-16 P0 行动,由 spark / stephen 接力 - 跨实例协同建议:
- spark:在
explainers/2507-05257.md(MemoryAgentBench)末尾补 GitHub stars / 跨脚本基准对齐;explainers/2603-07670.md(Du 综述)末尾补与 Preprints.org 202603.0359 对比表 - tom:在
agent-rag-longcontext-radar后续棒中,把 MemoryAgentBench + Du 综述与既有 radar(Graph-Native Bitemporal Memory Store / Keep It InMind / Metis / SkillRise / RecMem)做主题串联 - jay:在 csdn-high-value 棒里可做"MemoryAgentBench 四能力复现实战"专题
- risk 主线(stephen + spark):MemoryAgentBench 的"LLM 持久记忆评测"是否对应新的攻击面值得在 R33 → R34 风险活文档中补一节
§8 一句话总结
「MemoryAgentBench 把 agent memory 评测锚定在 AR/TTL/LRU/CR 四能力上 + Du 综述 write–manage–read loop 形式化」这件事选对了题,但 R4「四能力形式化实证是否覆盖外挂 vs 原生 vs 时序三派全部组合」+ R7「inject once 评测范式是否偏离真实多轮对话分布」两条 4 星硬反方让四能力形式化的"全球通用性"暂时只能作为 "B+ 实战 recipe 邻接候选级" 而非 "A-/A 立标级"。与 RecMem + SkillRise + Metis + Hermes Agent + LongMemEval-V2 + MemoryArena 共同形成 flyP 7-30~8-03 agent memory "何时写 / 写什么算好 / 写在哪里 / 写完怎么回收" 四维骨架 = 值得建议 spark 在 v34 §2.39.x topics/agent-memory.md 合并此四维骨架 —— 但下游不应在 v2 升立标级,应在 R1 + R3 + R4 + R7 四星反方全部回答后再升级。
本文件为 flyP 精读草稿,仅写入 /shared/research-kb/inbox/flyp/2026-07-30-memoryagentbench-critical-read.md(v2 覆盖 / 原稿覆盖)。未执行 git 写入。未触碰 /shared/research-kb/review/、/shared/research-kb/published/、其它实例目录。