Tom 文献雷达 · AI Agent / RAG / 长上下文 · 2026-07-23T20:40
本期候选(8 条,含 2 条新增)
| # | 来源 | 标题 | 标签 | 信号 |
|---|---|---|---|---|
| 1 | arXiv | IteraSim RAG: Multi-Stage Agentic RAG for OpenFOAM CFD | agent, rag, systems | 2026-07-22 |
| 2 | HF Daily | Scaling Laws for Hypernetwork-Based Knowledge Injection in LLMs | systems | votes:9 |
| 3 | HF Daily | Self Gradient Forcing: Native Long Video Extrapolation | multimodal | votes:25 |
| 4 | HF Daily | An Exam for Active Observers (ActiveVision) | benchmark, multimodal | votes:12 |
| 5 | HF Daily | FVAttn: Adaptive Sparse Attention for Video Generation | systems | votes:4 |
| 6 | HF Daily | Trace: Taxonomy-Guided Multidomain Visual Reasoning Env | multimodal | votes:3 |
| 7 | HF Daily | DocOps: Verifiable Benchmark for Autonomous Agents in Complex Doc Operations | agent, benchmark, systems | votes:2 |
| 8 | HF Daily | Train the Model, Not the Reader: Decodability Supervision | multimodal | votes:1 |
注:#1–#6 与 14:40 雷达重复,详见 2026-07-23T1440-agent-rag-longcontext-radar.md。
⭐ 高价值条目(2 条新增)
1. DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations ⭐⭐
- 来源: arXiv / HF Daily · 2026-07-21
- 地址: https://arxiv.org/abs/2607.19865
- 核心: DocOps 是一个层次化可验证评测框架,将真实世界文档操作(电子表格、表单、报告等)拆解为原子维度 + 递进工作流复杂度,对 autonomous agent 的文档可靠操控能力进行系统性评估。涵盖封闭和开源模型的多种 agentic harness 横向评测。
- 关键洞察: 该基准直接填补了"agent 操作真实数字文档"的可验证评测空白——不同于现有 benchmark 多聚焦于问答或代码,DocOps 评测的是 agent 能否可靠完成可验证的文档修改动作,对 AI 助手自动化办公工作流有直接参考价值。
- 标签: agent, benchmark, systems
2. Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations ⭐⭐
- 来源: arXiv / HF Daily · 2026-07-21
- 地址: https://arxiv.org/abs/2607.20379
- 核心: 研究当前"通过重建得分评价激活解释忠实度"的系统性缺陷:翻转个别错误claims后重建分数不变,说明模型在追踪具体事实而非整体要旨(gist)。提出 Decodability Supervision,在合成 ground truth 环境下对解释的具体claims进行逐条校验。
- 关键洞察: 对 agent 可解释性方向有反向启示:重建分数高不等于解释忠实,agents 用内部激活解释做决策时,不能只看重建得分,需逐 claim 验证。这对需要验证 agent 推理链忠实度的场景(法律、金融、医疗)有评测意义。
- 标签: multimodal, interpretability
📌 实战参考(来自 Substack)
Lesson 44: Evaluating Agentic RAG Reliability
来源:aiamastery.substack.com · 2026
- Ragas + Gemini 即时评测:每次推理同时产出可评估 artifacts,无需离线批量重跑
- Faithfulness 质量门禁:得分 1.0 = 答案每个 claim 均可回溯至检索段落;生产环境建议设质量底线而非追求满分
- MetricsEngine 双轨:Ragas 自动评分 + Gemini judge 覆盖 Ragas 暂无指标
- Live Dashboard:React + Recharts 追踪 per-question breakdown 和回归趋势
去重说明
- 本期 8 条候选中 6 条(#1–#6)已于 14:40 雷达覆盖
- #7 DocOps 和 #8 Train the Model, Not the Reader 为本期新增,首次进入本雷达
- 今日已产出 3 次雷达(08:40 / 14:40 / 20:40),arXiv 新量较低
Tom · 2026-07-23T20:40 · agent-rag-longcontext · 候选 8 条 · 高价值 2 条(新增)· Substack: 1 条 · 无 CSDN