Tom 文献雷达 · Agent × RAG × Long-Context · 2026-10-07 14:40 UTC

本期概览

来源:arXiv metadata + HF Daily;轻量快速综述。


🔴 高价值(3 条)

1. UNREAL: Unifying Retrieval and Long-Context with a Single Model

  • arXiv: 2610.08463 · 2026-10-06
  • 标签: rag long-context benchmark systems
  • 核心:用模型内部表征统一 RAG 检索(大规模语料)和长上下文推理;新增 < 500K 可训练参数,3B-token Wiki 实验。
  • 意义:RAG vs. Long-Context 架构分歧走向统一,值得追踪。

2. RAG-PIBench: Prompt-Injection Detection in Trustworthy RAG Systems

  • arXiv: 2610.08571 · 2026-10-06
  • 标签: rag benchmark systems
  • 核心:RAG 场景下提示注入检测基准,4,876 样本含 train/val/protected-test;DistilBERT F1=0.896。
  • 意义:RAG 安全评测填补空白,企业部署 RAG 必参考。

3. EC-RAG: Event Chain RAG for Long Video Understanding

  • arXiv: 2610.08674 · 2026-10-06
  • 标签: rag benchmark multimodal
  • 核心:将视频组织为事件链后再做 RAG QA,训练-free 框架,提升时序依赖建模。
  • 意义:RAG 在多模态长视频上的具体落地方式,非玩具演示。

🟡 中等价值(5 条)

4. HLA: Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

  • arXiv: 2610.05842 · 2026-10-05
  • 标签: long-context memory
  • 核心:HLA 为 Gated DeltaNet 引入查询相关 chunk 级混合,提升稀疏远程信息访问。

5. In-Parameter Memory Augmentation for LLMs(综述)

  • arXiv: 2610.08630 · 2026-10-05
  • 标签: agent memory systems
  • 核心:将知识编码进模型参数/适配器而非上下文;对比 ICL 和 Agent 方案的系统性综述。

6. Tool-Using Agents: Evidence to Action Failure Analysis

  • HF Daily · 2610.07753 · 20 votes · 2026-10-05
  • 标签: agent
  • 核心:10 个模型 harness 评测,静态评估强但交互执行弱;失败多发生在执行前(调查不完整或证据未确立就行动)。

7. MiniCorp: AI Agent Firm Simulator

  • HF Daily · 2610.05912 · 2026-10-04
  • 标签: agent
  • 核心:用电商模拟器生成企业级 agent 纵向数据;解决真实企业数据稀缺/隐私问题。

8. Taming VLAs: Self-Compensation under Robot Execution Errors

  • HF Daily · 2609.37334 · 15 votes · 2026-09-28
  • 标签: multimodal
  • 核心:VLA 在线自适应补偿执行误差,无需任务奖励或标签;跨执行条件鲁棒性压力测试。

💡 趋势洞察

  • RAG × Long-Context 统一:UNREAL 代表 architecture 收敛方向,而非两者对立。
  • RAG 安全成独立赛道:PIBench 标志 RAG 安全评测走向系统化。
  • Agent 评估受关注:Evidence-to-Action failure analysis 揭示"动作正确≠推理正确"的系统性偏差。
  • 记忆方向多元化:In-Parameter Memory 是继 ICL 之后的新范式,探索参数内化知识。

Caveat: Substack 本轮未检索到明确高价值新帖,主要依赖 arXiv/HF。