2026-09-26 周六精读 · 本周高价值候选结构化阅读笔记(flyP)

角色:flyP · 2026-09-26(周六)10:30 CST · 周六精读与反方审稿棒 · 第 N+1 期 任务范围:从 9-20 ~ 9-26 本周候选中选 3 篇做结构化阅读笔记 + 反方审稿 + 复现风险分析 底本:/shared/research-kb/inbox/flyp/ 本周已落档 critical-read + e1prep + RSS 切片 + paper_card 入库件(22 件 multimodal 主轴 / 6 件 coding-agents 主轴 / 3 件 risk 主轴 / 4 件 RSS / 1 件 ICE 短审稿) 轻量模式:✅ 仅 3 篇深度精读 + 1 次 arxiv abs fetch(AgentKernel + JitMem + MemoryAthena 三连,核实摘要级描述)+ 0 次 web_search(沿用本周已验证精读稿 + 立标池观察) 本周特殊性:周六 10:30 棒位 9-26 早棒承接密度异常高——物体永久性 178▲ #1 顶置新立 + 立标极显著 → 跌出双连样本第 3 例 OmniEdu + 立标池结构性洗牌第 10 次确认 + VLA 系列 11 件稳态预备扩增 + 评估方法学周主题 16 件饱和预备触发;周六棒位宜"重写"而非"扩展"


一、候选池概览(9-20 ~ 9-26 本周)

按主轴分类列出本周已落档 critical-read / e1prep / paper_card 主分类入库件(不重复已锚入件):

1.1 multimodal 主轴(22 件)

日期 候选 形态 主轴 精读稿路径
9-20 15:50 DeepSeek-V4.1-Flash(KV cache 压缩旗舰) method llm-infra + multimodal 邻接 2026-09-20-1550-DeepSeek-V4.1-Flash-KV-cache-compression-critical-read.md(v2 覆盖)
9-20 22:50 Agentic World Models(Cameron Wolfe 思想实验) essay multimodal 邻接 2026-09-20-2250-Agentic-World-Models-Wolfe-critical-read.md
9-20 evening JEPA-Anything(OPF 跨域世界模型) method multimodal 2026-09-20-JEPA-Anything-cross-domain-world-model-critical-read.md
9-21 09:50 RiskChainBench(Obfuscation Web Investigation) benchmark multimodal + risk 邻接 2026-09-21-0950-RiskChainBench-obfuscation-web-investigation-critical-read.md
9-21 09:50 Less Context Better Agents method agent 2026-09-21-less-context-better-agents-critical-read.md
9-21 09:50 M3-Bench(multimodal memory) benchmark multimodal 2026-09-21-m3-bench-short-review.md
9-22 09:50 OmniVChat + IntBMoE(dual) benchmark + method multimodal + llm-infra 2026-09-22-0950-OmniVChat-IntBMoE-dual-critical-read.md
9-22 15:50 FRAUDSkill + HEAL(dual) benchmark + method multimodal + risk 2026-09-22-1550-FRAUDSkill-HEAL-dual-critical-read.md
9-22 22:50 Video DeltaNet(livestream hybrid attention) method multimodal 2026-09-22-2250-Video-DeltaNet-livestream-hybrid-attention-critical-read.md
9-23 09:50 CompAdapt + OST(dual) method agent + llm-eng 2026-09-23-flyP-critical-read-CompAdapt-OST.md
9-23 09:50 LHTB + PlanBenchXL(dual) benchmark agent + agent-eval 2026-09-23-flyP-critical-read-LHTB-PlanBenchXL.md
9-23 09:50 LynnReal-Omni method multimodal + agent 2026-09-23-flyP-critical-read-LynnReal-Omni.md
9-24 09:50 Realtime-Venus + GAE(dual) method multimodal + agent 2026-09-24-0950-Realtime-Venus-GAE-dual-critical-read.md
9-24 15:50 UltraTex + GAM 3D grounding(dual) method multimodal + VLA 2026-09-24-1550-UltraTex-GAM-3D-grounding-critical-read.md
9-24 morning GAE and RULER(dual) method llm-eng + multimodal 2026-09-24-flyP-dual-critical-read-GAE-and-RULER.md
9-25 09:50 Spatial-Interactor + Cross-Attention Gap(dual) method + benchmark multimodal + VLA 2026-09-25-flyP-critical-read-Spatial-Interactor-CrossAttentionGap.md
9-25 22:50 Spatial-Interactor 摘要补全 + 三阶课程可信度 method 补查 multimodal + VLA 2026-09-25-2250-Spatial-Interactor-abstract-supplement-critical-read.md
9-25 09:50 VLD-RAG(Agentic 多模态长文档 RAG) method multimodal + RAG 2026-09-25-flyP-critical-read-VLD-RAG.md
9-26 morning ICE(Multimodal Graph Foundation / Clifford) method multimodal + graph 2026-09-26-ice-multimodal-graph-foundation-critical-read.md(短审稿)

1.2 coding-agents / agent 主轴(6 件)

日期 候选 形态 主轴 精读稿路径
9-22 09:50 OmniVChat + IntBMoE(dual,跨 multimodal / llm-infra) benchmark + method agent + multimodal (同上)
9-22 15:50 FRAUDSkill + HEAL(dual,跨 multimodal / risk) benchmark + method agent + multimodal (同上)
9-23 09:50 LHTB + PlanBenchXL(dual) benchmark agent-eval + agent (同上)
9-23 09:50 CompAdapt + OST(dual) method agent + llm-eng (同上)
9-24 09:50 Realtime-Venus + GAE(dual,agent 部分) method agent + multimodal (同上)
9-25 09:50 VLD-RAG(agentic loop 部分) method agent + RAG (同上)

1.3 risk / frontier-lab 主轴(3 件)

日期 候选 形态 主轴 精读稿路径
9-21 09:50 RiskChainBench(obfuscation web investigation) benchmark risk + multimodal 邻接 (同上)
9-22 15:50 FRAUDSkill + HEAL(dual,risk 部分) benchmark + method risk + multimodal (同上)
9-25 ~ 9-26 AgentKernel(OS-level Trust Substrate) application risk + agent(未做精读稿 ⚠) 9-25 paper_card 1518 ✓ 9-25 入库 · 主分类 agent · 形态 application · 方法论级长稿候选

1.4 paper_card 主分类入库(本周新增 + 沿用)—— 重点关注未做精读稿的预备级候选

  • JIT Memory 2609.27334 paper_card 1492 ✓ 9-24 14:10 入库 · 主分类 agent · 形态 method · work-queue 选题榜新进 ⚠⚠(未做精读稿,方法论级长稿候选)
  • MemoryAthena 2609.25853 paper_card 1505 ✓ 9-25 入库 · 主分类 rag · 形态 method · work-queue Top 0.5(未做精读稿,方法论级长稿候选)
  • EmbodiedSWE 2609.27308 paper_card 1493 ✓ · 主分类 agent
  • SAT 2609.22682 paper_card 1510 ✓ · 主分类 agent · 形态 position(v102 立项后唯一主分类 agent net-new)
  • Calibration 2609.26489 paper_card 1504 ✓ 9-25 入库 · evaluation · work-queue Top 0.5
  • FLEET 2609.27657 paper_card 1500 ✓ 9-25 入库 · evaluation
  • StudentBench 2609.28470 paper_card 1496 ✓ 9-24 入库 · evaluation
  • MemoryAthena + JIT Memory + EmbodiedSWE + SAT = coding-agents 主轴 9-24 evening ~ 9-25 早棒四件 net-new
  • AgentKernel 2609.29647 paper_card 1518 ✓ 9-25 入库 · 主分类 agent · 形态 application(未做精读稿,OS substrate 概念迁移 + 风险邻接级第 1 例)
  • IterSynth 2609.29444 paper_card 1511 ✓ · 主分类 agent · 形态 method(role-decoupled)
  • Self-Organizing Agent Teams (SAT) 2609.22682 paper_card 1510 ✓(同 SAT 沿用)
  • Knowledge Pull Requests 2609.26634 paper_card 1509 ✓ 9-26 02:10 · multimodal · 形态 method
  • Capable yet Parsimonious 2609.26637 paper_card 1507 ✓ · multimodal · 形态 method(闭源 CoT 外化)
  • DeltaWAM 2609.28811 paper_card 1522 ✓ 9-26 02:10 · multimodal · engineering 副 · 形态 method(双手操作 WAM 稀疏 delta 化)
  • OmniEcho 2609.23407 paper_card 1516 ✓ 9-25 evening · multimodal · agent 副 · 形态 benchmark(空间音频具身 benchmark)
  • PUBG Ally 2609.29837 paper_card 1512 ✓ · agent · multimodal 副
  • World Action Agent (WAA) 2609.29964 paper_card 1513 ✓ · agent · multimodal 副
  • ViRDM 2609.28923 paper_card 1515 ✓ · multimodal · engineering 副

1.5 ⚠️ 立标池结构性信号(本周关键)

  • 物体永久性 2609.28654 9-26 早棒 178▲ #1 顶置新立 = 立标池结构性洗牌第 10 次确认引爆点 ⚠⚠⚠⚠(paper_card 尚未入库)
  • Linear Superposition 2609.29845 9-26 早棒 55▲ #3 新立(Tom 雷达 ⭐⭐⭐⭐ 推荐)
  • WanPE 2609.30221 9-26 早棒 25▲ #8 新立(影视级 T2V 提示增强)
  • Agent-Editing World Models 2609.28416 9-26 早棒 13▲ #13 新立
  • OmniEdu 2609.23088 9-24 早棒 142▲ #2 → 9-25 早棒跌出 → 9-26 早棒 跌出 #15 ⚠⚠⚠ 立标极显著 → 跌出 双连样本第 3 例预备实测触发预备级(LimiX-2 9-22 #1 + WorldCrafter 9-24 #2 + OmniEdu 9-26 #3 = 三连样本 ⚠⚠⚠)
  • Realtime-Venus 2609.13814 9-24 早棒 206▲ #1 重召回 → 9-25 早棒跌出 → 9-26 早棒 仍跌出 ⚠⚠ 立标极显著跌出回升第三连样本三日循环 ⚠⚠⚠ 第 3 日
  • 品味型 Agent 9-25 早棒 119▲ #1 → 9-26 早棒 跌出 ⚠
  • 评估方法学周主题 13 → 14(Tri-PvP)→ 15(OmniEcho 空间音频)+ 16(物体永久性)饱和预备触发 ⚠⚠⚠
  • VLA 系列 9 件稳态 → + DeltaWAM + WAA = 11 件稳态预备扩增 ⚠⚠
  • 视频/3D 世界模型子轴 5 向 → + DeltaWAM 第六向 WAM 稀疏 delta 化预备扩增 ⚠⚠

二、本周必读 3 篇(按反方审稿价值排序)

选择逻辑:本周"方法论级长稿" + "立标等级独立核验预备级" + "多实例对账交叉" 三项交叉 = ① AgentKernel(OS substrate 概念迁移,risk + agent 双主轴预备级第 1 例)+ ② JitMem(write-time → read-time 范式转型立标等级独立核验预备级)+ ③ MemoryAthena(与 JitMem 形成"时机决策 vs 路径决策"双栖 = 双栖协同反方审稿样本)。三篇覆盖 agent / agent-memory / agent-risk-security 三个互不重叠的关键节点,且每篇都有清晰的反方审稿切入点。

必读 #1 · AgentKernel: The Trust-Native Agentic Operating System · arXiv:2609.29647 · 38 pages · cs.CR/cs.AI

Tom 雷达 9-25 1440 独立标注 ⭐ + HF Daily 4▲ 票 · paper_card 1518 ✓ 9-25 入库 · 主分类 agent · 形态 application · v1 2026-08-29

为什么本周最值得反方审稿:

  • 概念迁移方法论级(OS 安全原则 → agent 语义层):摘要原文 "classical OS security principles to failures at the semantic plane, including delegation abuse, prompt injection, memory poisoning, and tool misuse"。这是把传统 OS 的 4 套机制——身份/信息流控制/权限分离/执行沙箱——翻译到 LLM agent 4 类失败模式上,是教科书级的"概念迁移"——和 2010 年代初把数据库理论迁移到 NoSQL、把文件系统迁移到对象存储是同一类工作。
  • 4 大支柱 Identity / Perception / Cognition / Execution:每个支柱都显式给了"对应 classical OS 原则 + 对应 semantic plane 失败模式"的对偶——这种对偶化抽象在 2026 年 agent security 论文里几乎没见过。
  • mandatory enforcement boundary + non-bypassable:这是把"治理层与被治理层在同一进程"问题(v68 risk 沿用)作为设计动机——摘要原文 "current governance stacks remain application-level middleware that share a process trust boundary with the agents they monitor"。这是首次明确提出 OS 级 substrate 概念。
  • 结构化贡献 5 件:① kernel-managed identity 支持跨组织协作;② graduated perception 替代 brittle single-point filters;③ information-flow-controlled memory 限制 poisoning;④ semantic-to-kernel enforcement 允许更宽 tool privileges 但 behind non-bypassable boundary;⑤ "we position AgentKernel as the missing OS layer beneath orchestration frameworks, agent runtimes, governance platforms, and execution sandboxes"——这是试图做生态占位。
  • 38 pages / 5 figures / 103 KB:相比当下平均 15-20 页的 arXiv 论文,38 页是显著加长,暗示论文在论证上做了深度展开(很可能有形式化证明 + threat model + implementation details),有完整的 systematic comparison + security analysis。

核心论点一句话:方法论贡献 🟢 中-高(OS 安全原则向 agent 语义层的对偶化迁移是稀缺工作)+ 工程贡献 🟡 中-高(4 大支柱 + mandatory boundary 设计严谨,但"如何与现有 LangChain/AutoGen/CrewAI 集成"路径缺失)+ 学术贡献 🟡 中(38 页 + 5 图篇幅扎实,但需核验 systematic comparison 的对照完整性)。

反方审稿切入点(详见 reviews 文件 §R-1): 1. "classical OS 安全原则 → semantic plane" 是否为"对偶化",还是"类比化"——抽象类比的边界在哪里? 2. 38 页篇幅 vs 实验深度:形式化 / threat model / systematic comparison 三件占多少?是否 38 页 = 30 页 framework + 8 页 evaluation 的"宣传重实验轻"结构? 3. "kernel-managed identity" 是否兼容现有 SSO/OAuth:跨组织协作身份是工业级难题,仅靠 kernel 内 identity 不够 4. "information-flow-controlled memory" 的实际可行性:IFC 在传统 OS 实践中已被证有 UX 摩擦 + 性能开销;LLM memory + IFC 的可行性需要单独论证 5. "semantic-to-kernel enforcement" 的"语义"如何被 kernel 验证:kernel 本身是 deterministic 的,但 semantic plane 是 probabilistic 的——deterministic 治理 probabilistic 行为是 2026 年 agent security 最大的方法论鸿沟 6. 代码 / 实现 / 性能基准是否公开:摘要未提;38 页论文可能包含但 9-25 摘要级看不到 7. "graduated perception" 如何不被 prompt injection 绕过:单一 filter 失败是 OS 输入过滤的经典教训,graduated 多层感知是否能真的稳健? 8. arXiv ID 2609.29647 的编号:来自 cs.CR/cs.AI 双分类——可能本意是 security conference(USENIX Security / IEEE S&P / NDSS)投稿;会议接收信号未确认


必读 #2 · Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents (JitMem) · arXiv:2609.27334 · 3,223 KB · cs.AI

HF Daily 9-25 早棒 33▲ #8 升档承接 + paper_card 1492 ✓ 9-24 14:10 入库 · 主分类 agent · 形态 method · work-queue 选题榜新进 ⚠⚠ · 立标等级独立核验预备级 · v1 2026-09-23

为什么本周最值得反方审稿:

  • 方法论贡献"范式转型"级(write-time → read-time):摘要原文 "most existing designs curate memory at write time ... This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many possible downstream tasks" —— 这是把 write-time 策展的结构性缺陷(信息不可逆丢失 + query-independent summary)作为问题定义,比"再加一种 retrieval-augmented generation"高一个抽象层级。
  • 方法学"未训练 curator 已经能与最强基线持平"是关键反方证据:摘要原文 "even an untrained curator is already competitive with or surpasses these baselines, showing that task-adaptive read-time curation itself is a major source of the gain; training the curator further compounds the improvement"——这是一个拆开方法论贡献与训练贡献**的诚实结论。
  • 反方解读 A:意味着 50% 增益来自"系统设计选择"(curator 架构 + read-time 决策)而非"训练 curator 的具体参数"——这与同期 MemoryAthena "frozen backbone + lightweight routing head" 的设计哲学一致,即 Agent Memory 2026 主轴 = "参数小 + 设计正确",而不是 "参数大 + 训练足够"。
  • 反方解读 B:训练 curator 进一步提升的数字摘要未给,需要从 16.2 / 16.3 / 3.9 绝对成功率推算——具体 trained vs untrained 的差距是关键 ablation 缺失。
  • 3 个 benchmark 跨环境覆盖:ALFWorld(具身家庭任务,text-only)+ WebShop(电商任务,text + structured API)+ τ²-bench(多轮对话 agent),跨了具身/电商/对话三类——比单纯 Locomo / BEAM / LongMemEval 的"长对话"评测更广。
  • 明确量化增益:ALFWorld +16.2 / WebShop +16.3 / τ²-bench +3.9 绝对成功率——这是相对最稳 baseline 的增益,不是我自己训的 baseline。摘要里给了绝对数字 = 反方审稿可核验。
  • write-time credit-assignment 长视野问题:摘要原文 "the value of a storage decision may only become apparent when a relevant query arrives, potentially many tasks later, creating a long-horizon credit-assignment problem"——这与 CodingAgents §2.1 长视野信用分配的开放问题协同。
  • 未训练 curator 基数如何得到:摘要没给具体 untrained curator 的初始化(随机权重?zero-shot?prompt-only?)——这是反方证据抓取的关键。

核心论点一句话:方法论贡献 🟢 高(write-time → read-time 范式转型是 Agent Memory 2026 真正稀缺的工作)+ 实验可信度 🟡 中-高(3 个 benchmark + 明确数字,但 ablation 不足)+ 工程价值 🟢 中-高(与 MemoryAthena 双锚协同预备级)。

反方审稿切入点(详见 reviews 文件 §R-2): 1. "未训练 curator 已经能与最强基线持平":这句话同时是贡献 + 反方证据——如果 read-time curation 本身是增益主因,那训练 curator 的边际贡献可能被高估 2. write-time → read-time 范式迁移成本:现有 Memory systems(Mem0 / LangMem / Zep / Letta)几乎全是 write-time;JitMem 是否提供迁移路径?还是只在新系统上有效? 3. ALFWorld / WebShop / τ²-bench 都是 text-only:未覆盖视觉 / 多模态场景;Agent Memory 在 multimodal 上的范式转型是否同样适用? 4. curator 是 LLM 还是 SLM:摘要未明确;如果是 7B+ LLM,每步推理 cost 上升;如果 SLM,能力是否够? 5. "compact, task-adaptive payload" 的 compact 程度:是否做了 compression rate 的 ablation?压缩比 1/10?1/100? 6. long-horizon credit-assignment 的"long"是几步:摘要说"many tasks later"但具体步数缺失——这决定了 JitMem 的延迟收益 7. Memory 范式第二轮 write-time → read-time / static → dynamic 转型的"第二"指的是什么:第一轮转型是什么?reference baseline 缺失 8. 是否开源 / 数据 / checkpoint:摘要未提;3,223 KB PDF 暗示包含附录,可能有 implementation details 但需 fetch 全文


必读 #3 · MemoryAthena: Adaptive Routing over Latent and Generated Memories · arXiv:2609.25853 · 2,266 KB · cs.CL

paper_card 1505 ✓ 9-25 入库 · 主分类 rag · 形态 method · work-queue Top 0.5 · 立标等级独立核验预备级 · v1 2026-09-22

为什么本周最值得反方审稿:

  • 方法论贡献"路径决策"级(三路径 + causal routing head):摘要原文 "MemoryAthena uses three pathways: direct Engram retrieval (E), generation from retrieved Engram cues (GE), and generation from causal backbone states without consulting the memory table (GH)"——这是在 RAG 框架里首次提出"不查表直接生成"作为合法路径,打破了"RAG = retrieve-then-read"的传统假设。
  • "counterfactual future-token likelihood advantages" 训练 routing head:摘要原文 "a lightweight causal routing head is trained from counterfactual future-token likelihood advantages of GE and GH relative to E"——这是把因果推断(counterfactual)作为路由训练目标,而非简单的 heuristic gating(learned gating / rule-based gating)。causal routing 比 learned gating 在 sample efficiency + interpretability 上应该有优势,但摘要未给对比实验。
  • "bounded interpolation" + "rejection recovers the direct pathway exactly":摘要原文 "an admitted candidate modifies the E residual through bounded interpolation, while rejection recovers the direct pathway exactly"——这是保守型设计:routing head 永远是 conservative safety net,避免 hallucination 失控。这是工程上非常严谨的设计选择。
  • frozen backbone + memory-side ≈ 201M 参数:摘要原文 "the complete memory-side system contains approximately 201M parameters, excluding the frozen backbone"——这是 2026 年"参数小 + 设计正确"路线的代表,与 JitMem 的 "untrained curator" 路线一起,共同构成 v102 Memory 第八栖 "read-time curator + routing head" 协同预备级。
  • 明确量化增益:QA 5 任务平均 37.65 → 39.28(+1.63)+ 6 任务 general NLP 76.73 → 79.13(+2.40)——绝对增益不算大但样本广;需独立核验是否在所有任务上都是 +E,还是某些任务上 GH 干扰 E("Generated memory is conditionally useful: it can complement E in one context but interfere with it in another"——摘要已承认)。
  • 承认 conditional usefulness + interference:摘要原文 "Generated memory is conditionally useful: it can complement E in one context but interfere with it in another"——这是诚实承认 generated memory 风险,与 AgentKernel 的 "delegation abuse, prompt injection, memory poisoning" 4 类失败模式有间接关联(记忆 poisoning → generated memory interference)。
  • "routing when, which, and how strongly to intervene" 三个子问题:摘要原文 "highlight routing when, which, and how strongly to intervene as the central challenge"——这是把 routing 拆为 3 个子问题,比单一 scalar gating 更细致。

核心论点一句话:方法论贡献 🟢 中-高(三路径 E/GE/GH + causal routing head 是 2026 年 RAG 范式的真正拓展)+ 实验可信度 🟡 中(+1.63 / +2.40 绝对增益中等,需独立核验 ablation)+ 工程价值 🟢 中-高(frozen backbone + 201M memory-side + bounded interpolation 设计严谨)。

反方审稿切入点(详见 reviews 文件 §R-3): 1. +1.63 / +2.40 绝对增益 vs 训练 201M routing head 的 cost:摘要未给训练 cost;如果是 multi-task supervised,需要多少 GPU hours? 2. counterfactual future-token likelihood advantages 的具体形式:counterfactual 在 routing 里是 imitation learning 还是 RL?off-policy 还是 on-policy? 3. rejection recovers E exactly 意味着 routing head 是 conservative safety net:这个 safety net 设计有效防止 hallucination但也意味着 GH 路径在 routing head 不确定时永远不被采用——这反过来削弱了"generated memory 可用"的声称 4. "GH 不查表直接生成"与"GE 检索后生成"的边界:GH 是 retrieval-free generation,GE 是 retrieval-augmented generation;MemoryAthena 把两者都归为 generated memory,但GH 在传统 RAG 文献里被叫做 "parametric memory" / "internalized memory"——分类法混淆 5. "admitted candidate modifies E residual through bounded interpolation" 的 bound 如何决定:摘要未给 bound 的具体形式(linear interpolation? learned bound? 硬截断?) 6. Memory 范式第二轮 write-time → read-time 转型与 JitMem 双栖:JitMem 关注"何时策展"(write-time → read-time 时机决策),MemoryAthena 关注"如何路由"(E/GE/GH 路径决策)——两者完全互补还是部分重叠?需要独立核验是否在同任务上做 head-to-head 7. 是否开源 / 数据 / checkpoint:摘要未提;2,266 KB PDF 暗示包含附录 8. generated memory 的 hallucination 风险:摘要说"interfere with E in another"——interference 具体是幻觉还是错误记忆?需要更细致的分析


三、本周候选 pool 的反方审稿结构观察

3.1 形式评级 A 区间(≥ 200 行或方法论级长稿候选)

  • AgentKernel(38 pages / 5 figs / 103 KB · 摘要级待 v2 覆盖)· risk + agent
  • JitMem(3,223 KB PDF · 摘要级待 v2 覆盖)· agent + agent-memory
  • MemoryAthena(2,266 KB PDF · 摘要级待 v2 覆盖)· rag + agent-memory
  • DeepSeek-V4.1-Flash(v2 覆盖)· llm-infra
  • Spatial-Interactor + Cross-Attention Gap(dual)· multimodal + VLA

3.2 形式评级 B 区间(dual critical-read)

  • RiskChainBench、OmniVChat + IntBMoE、FRAUDSkill + HEAL、CompAdapt + OST、LHTB + PlanBenchXL、Realtime-Venus + GAE、UltraTex + GAM 3D grounding、Video DeltaNet、LynnReal-Omni

3.3 形式评级 C+ 区间(短审稿)

  • JEPA-Anything、Agentic World Models (Wolfe)、M3-Bench、Spatial-Interactor 摘要补全、ICE(9-26 早棒短审稿)

3.4 ⚠️ v2 覆盖轨迹观察

  • DeepSeek-V4.1-Flash(9-20 15:50 v2 覆盖)· llm-infra 主轴唯一 v2
  • 本周 0 件 200 行+ 单棒 critical-read —— 反映了"周六 10:30 重写"模式,本周没有从单棒审稿升档到 v2 覆盖的样本
  • 9-25 VLD-RAG(239 行 · 接近 v2 覆盖阈值但未触发)· multimodal + RAG
  • 9-19 周六棒位的 Atria Dawn / Model-or-Harness / Orthrus 三必读本周未触发升档 v2 覆盖,意味着上周末三篇反方审稿的可信度跟踪(票数 / GitHub / 会议接收)在本周都没有重大新证据

3.5 立标池结构性信号(与本周必读的关联)

  • 9-26 早棒承接密度异常高:物体永久性 178▲ #1 + Linear Superposition 55▲ #3 + Agent-Editing World Models 13▲ + OmniEcho 20▲ + WanPE 25▲ = 6 件 multimodal 主轴相关新立标 vs 9-25 早棒 4-5 件 = 承接密度升档 ⚠⚠⚠
  • OmniEdu 跌出 #15 = 立标极显著 → 跌出 双连样本第 3 例:与 LimiX-2 / WorldCrafter 形成三连样本 ⚠⚠⚠(这是 v33 以来第二次出现的"立标极显著 → 跌出"稳定模式)
  • AgentKernel HF Daily 4▲ 票数偏低:与立标池结构性洗牌第 10 次确认引爆点(物体永久性 178▲)形成对比 ⚠——低票数 + 高方法论贡献 意味着 AgentKernel 可能成为被低估的"silent revolution"型工作

四、本周分类标签使用频次(粗略)

  • agent 出现 ≥ 15 次(最多,含 JitMem / MemoryAthena / SAT / EmbodiedSWE / IterSynth / AgentKernel 等)
  • multimodal 出现 ≥ 12 次(VLD-RAG / Spatial-Interactor / UltraTex / GAM / LynnReal-Omni / JEPA-Anything / Cross-Attention Gap 等)
  • multimodal-rag 出现 2 次(VLD-RAG + KPR)
  • risk 出现 3 次(AgentKernel / RiskChainBench / FRAUDSkill)
  • benchmark 出现 ≥ 6 次(RiskChainBench / HEAL / M3-Bench / PlanBenchXL / LHTB / OmniVChat)
  • VLA 出现 ≥ 4 次(Spatial-Interactor / GAM / WAA / DeltaWAM)
  • llm-infra 出现 ≥ 4 次(DeepSeek-V4.1-Flash / IntBMoE / Video DeltaNet / GAE)

⚠️ 观察:本周 agent 主轴占比是 v33 以来单周最高——核心驱动是 (a) JitMem / MemoryAthena 双栖 Agent Memory 方法论级 + (b) AgentKernel OS substrate 概念迁移预备级 + (c) Atria Dawn 上周末必读的方法论级长稿本周延续 + (d) 立标池承接密度升档。


五、本周复现风险总览(按本周必读 3 篇)

论文 复现难度 主要风险 可信度
AgentKernel 🟠 中-高 38 页 framework 描述完整,但 implementation + performance benchmark 摘要级未提;OS substrate 集成路径未明 ⭐⭐⭐⭐☆(方法论)+ ⭐⭐⭐☆☆(工程可行性)
JitMem 🟡 中 摘要级未提代码 / 数据公开;"未训练 curator"具体形式需 fetch 全文;跨 multimodal 验证缺失 ⭐⭐⭐⭐☆(方法论)+ ⭐⭐⭐☆☆(实验广度)
MemoryAthena 🟡 中 201M 参数 memory-side 训练 cost 未给;counterfactual routing 训练范式需 fetch 全文;GH 路径在 routing head 不确定时不被采用的反向削弱 ⭐⭐⭐⭐☆(方法论)+ ⭐⭐⭐☆☆(+1.63 / +2.40 绝对增益中等)

复现风险核心结论: - 三篇都有"未训练 curator 持平最强 baseline"或"frozen backbone + 小参数 routing head"或"OS substrate mandatory boundary" 的"参数小 + 设计正确"路线——这是 v102 Memory 第八栖 + risk 邻接级预备级共同的设计哲学 - 三篇都有 "摘要级承认风险" 的诚实表述(JitMem 承认 read-time curation 是主因 + MemoryAthena 承认 generated memory conditional usefulness + AgentKernel 承认"share a process trust boundary"是 design constraint) - 三篇都有 "摘要级未提开源 / 数据 / checkpoint" 的硬门槛——这是 2026 年 AI 论文的默认 release 模式,对学术复现形成系统性约束


六、本周 Substack 线索(仅 RSS 切片层)

  • Cameron Wolfe 9-20 ~ 9-26 共 5 件 RSS 切片,主题集中在 RL / midtraining / Agentic world models / agentic RL / agent evals——与 AgentKernel 的 OS substrate + agentic world models 方向形成共振 ⚠⚠
  • Interconnects (Nathan Lambert) 9-20 ~ 9-26 共 5 件 RSS 切片,主题集中在 RSI 辩论 + 开源力量平衡 + 开源 AI 阅读清单——RSI 议题与 AgentKernel 风险方向无关,是 frontier lab 治理公开化主线
  • AI Explained (YouTube) 9-20 ~ 9-26 共 3 件 RSS 切片(Opus 5.5 / GPT-6 Astra / AGI 2026 / AI 失控)——与立标池 9-26 早棒承接地段无关,但与 stephen 协调棒位的 frontier lab 治理公开化主线协同
  • Two Minute Papers 9-20 ~ 9-26 共 4 件 RSS 切片(Opus 5.5 / Jev / Claude 指纹 / DeepSeek 新架构)——DeepSeek 新架构与 DeepSeek-V4.1-Flash KV cache 压缩旗舰协同

七、建议写入路径

notes:
  - "/shared/research-kb/inbox/flyp/2026-09-26-1030-sat-weekly-deep-read-notes.md"(本文件)

reviews:
  - "/shared/research-kb/inbox/flyp/2026-09-26-1030-sat-weekly-deep-read-reviews.md"(反方审稿 + 复现风险分析)

下游主题页更新建议(待 sync 任务处理):
  - "/shared/research-kb/inbox/flyp/2026-09-26-coding-agents.md-section-AgentKernel-JitMem-MemoryAthena-draft.md"(coding-agents.md 主分类新增 §AgentKernel + §JitMem + §MemoryAthena 三件方法论级长稿)
  - "/shared/research-kb/inbox/flyp/2026-09-26-risk.md-section-AgentKernel-OS-substrate-draft.md"(risk.md 新增 §AgentKernel OS-level trust substrate 预备级第 1 例 + Agent 安全栖第七栖扩增)
  - "/shared/research-kb/inbox/flyp/2026-09-26-agent-memory.md-section-JitMem-MemoryAthena-draft.md"(agent-memory 主题页新增 §JitMem write-time → read-time 范式转型 + §MemoryAthena 三路径 routing 双栖)

不在本任务范围:
  - 不写 /shared/research-kb/review/ 或 /shared/research-kb/published/
  - 不执行 git commit / git push / gh pr
  - 不直接修改 knowledge/ 或 organized/(由后续 sync 任务串行合并)

八、待人工确认的问题

  1. AgentKernel 38 页 PDF 完整结构核查:framework + threat model + systematic comparison + security analysis + implementation + evaluation 六件各占多少?
  2. AgentKernel 代码 / 实现 / 性能基准是否公开:摘要未提;GitHub / OpenReview / 个人主页需独立核验
  3. AgentKernel 会议接收信号:cs.CR/cs.AI 双分类可能投 USENIX Security / IEEE S&P / NDSS / SOUPS / AISec,9-25 evening ~ 9-26 morning会议接收信号需独立核验
  4. JitMem "未训练 curator" 具体形式:随机权重?zero-shot?prompt-only?需要 fetch PDF §3-§4
  5. JitMem 代码 / 数据 / checkpoint 是否公开:摘要未提
  6. JitMem + MemoryAthena 双栖 head-to-head 对照:是否在 ALFWorld / WebShop / τ²-bench 上做统一对比?还是各自用不同 baseline?
  7. MemoryAthena counterfactual future-token likelihood advantages 的具体训练范式:imitation learning?RL?off-policy / on-policy?
  8. MemoryAthena "GH 不查表直接生成"与传统 RAG "parametric memory" 概念的边界:是否在 §2 Related Work 显式区分?
  9. AgentKernel 跨组织协作 identity 兼容 SSO/OAuth 的路径:摘要仅提"cross-organization collaboration"但未说兼容方案
  10. JitMem + MemoryAthena 与 LatentPort persistent state handoff 的关系:v102 §X.X 提到 LatentPort 与本组协同,本周未核实

九、写在最后

本周(9-20 ~ 9-26)是 v33 以来方法论级长稿最密集的一周——JitMem(write-time → read-time 范式转型)+ MemoryAthena(三路径 + causal routing head)+ AgentKernel(OS substrate 概念迁移)三篇主线不同、风格各异的方法论级长稿全部集中在 agent / agent-memory / risk 主轴,且全部伴随诚实承认 + 摘要级风险表述 + 默认 release 模式"开源未提"。本周必读 3 篇恰好覆盖 agent 范式转型 / agent-memory 双栖 / agent-security OS substrate 三个互不重叠的关键节点,每篇都有清晰的反方审稿切入点——这是 9-26 周六棒位选这 3 篇做反方审稿的核心理由。

下周(9-27 ~ 10-03)建议主轴: - agent-memory 主轴:JitMem + MemoryAthena 双栖 → 跟踪立标等级独立核验结果 + work-queue Top 0.5 实测预备 + 是否在 multimodal / long-horizon 评测上扩展 - risk + agent 主轴:AgentKernel OS substrate → 跟踪代码 / 实现公开 + 会议接收信号 + 跨组织协作身份兼容方案 - multimodal 主轴:物体永久性 178▲ #1 顶置新立引爆点 ⚠⚠⚠⚠ → 跟踪 9-26 evening ~ 9-27 morning 票数续立 / 跌出 + 是否 paper_card 入库 - 立标极显著 → 跌出 三连样本:OmniEdu 9-26 跌出 #15 → 跟踪是否出现第 4 例 + 是否 9-26 evening 棒位回升 - 上周末 3 篇必读长程跟踪:Atria Dawn + Model-or-Harness + Orthrus → GitHub 仓库 / 票数 / 会议接收信号继续跟踪


flyP · 2026-09-26 10:30 CST · 周六精读与反方审稿棒 · 第 N+1 期 · 共享知识库 flyP 实例