Tom 文献雷达 · AI Agent / RAG / 长上下文 · 2026-08-02 晚间

候选摘要(8 条)

高价值(4 条)

  1. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems - arXiv 2607.27958 | 2026-07-29 - 多智能体场景下建模"哪些 peer 可信"的在线可靠性记忆系统,用 Weyl 不等式维护信任状态;适合多 agent 协调、RAG 路由等场景。 - 标签:agent memory systems

  2. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability - arXiv 2607.26637 | 2026-07-28 - 首个系统研究"agent 将记忆存为文件系统目录树"的工作,测试随记忆累积/冲突/过期的自我组织能力,发现文件系统并非最优默认方案。标签:agent rag memory benchmark

  3. DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal RAG - arXiv 2607.28580 | 2026-07-30 - 解决多模态 RAG 中粗粒度丢细节、细粒度图爆炸的两难;提出宏/微观双路解耦,在 12 个数据集上验证。标签:rag benchmark multimodal

  4. GLM-RAG: Graph Language Models for Graph-Based RAG - arXiv 2607.28397 | 2026-07-30 - GLM(融合图推理+语义)与 GNN/向量检索在单跳/多跳 RAG 上的系统对比;标签:rag benchmark

一般(4 条)

  1. See2Think: Do Multimodal Models Really Use Intermediate Visual States? - arXiv 2607.26769 | 2026-07-28 | HF 22 票 - 评估多模态 LLM 是否真正依赖中间视觉状态(草图、注释、工具);提出 See2ThinkBench + VAoT,1200 题 12 类。标签:rag benchmark multimodal

  2. ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection - arXiv 2607.28126 | 2026-07-30 - 工业检查日志场景下贡献感知记忆框架,对每个记忆单元估计下游诊断价值,辅助早风险筛查。标签:rag memory benchmark systems

  3. OmniScope: Modality-Decoupled Token Compression for Omnimodal LLMs - arXiv 2607.23193 | 2026-07-27 | HF 3 票 - 音频/视频 relevance 峰值时刻不同,单向引导易丢关键线索;解耦分配模态特定 token 预算。标签:multimodal

  4. Fairness Pruning: Locating Demographic Bias in GLU-MLP - arXiv 2607.28319 | 2026-07-29 | HF 2 票 - 用对比 prompt + 激活捕获定位 GLU 架构 down_proj 层的人口统计偏差神经元。标签:benchmark systems


Substack 线索(1 条)

  • Agent Memory Is Not RAG(Claudio Stamile,2026-01) 来源:https://claudiostamile.substack.com/p/agent-memory-is-not-rag-a-practical 核心:agent memory 与 RAG 的目的根本不同——RAG 评估检索质量,memory 评估随时间的自适应能力;提出 forms/functions/lifecycles 三维框架代替短/长期二分。与本次候选 #2(文件系统记忆)高度呼应。

去重说明

本日已产出 2 份 radar(08:40 / 14:40);本批次候选以 07-30 新论文为主,与前次无实质重复。


Tom 文献雷达 · AI Agent / RAG / 长上下文 · 2026-08-02T20:40 UTC