📡 Tom 文献雷达 · Agent RAG LongContext · 2026-08-29 20:40
主题: Agent / RAG / 检索 / 长上下文 / 评测 / 新论文
来源: HF Daily (arXiv) + Substack
🔬 本期候选(8 条)
| # | 标题 | 来源 | 标签 | 信号 |
|---|---|---|---|---|
| 1 | PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents | arXiv 2608.26530 | agent, long-context | 26票 |
| 2 | CritICL: Inference-Time Weak-to-Strong Generalization from Small LM Failure Modes | arXiv 2608.27455 | rag, systems | 6票 |
| 3 | What Does an Evaluation License? A Commit-Bound Census of Inspect Evals | arXiv 2608.19269 | benchmark, systems | 2票 |
| 4 | Agentic Game Dev as Verifiable Trajectory Data Engine for Scaling World Models | arXiv 2608.25518 | agent, multimodal | 132票 |
| 5 | Procedura: Agentic 3D Modeling with Procedural Control | arXiv 2608.26238 | agent, systems | 9票 |
| 6 | Luce: Relightable Gaussians for 3D Asset Generation | arXiv 2608.23943 | multimodal | 7票 |
| 7 | TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback | arXiv 2608.25798 | multimodal | 5票 |
| 8 | EditaLive! Unified Character Video Editing for Live Streaming | arXiv 2608.27123 | multimodal | 3票 |
⭐ 高价值条目(3 条)
1. PILOT in the Loop — 实时自我改进长程 Agent
arXiv: https://arxiv.org/abs/2608.26530
现有自改进方法只能在任务执行结束后处理经验,无法中途重定向当前运行。PILOT 提出在线(live)自改进框架,在 agent 运行过程中同时:(a) 用 emerging experience 重定向当前任务执行,(b) 更新持久化 harness。核心洞见:将经验既用于当前 run 的实时调整,也用于未来任务的持续改进,而非等到运行结束再事后分析。
关联主题: Agent 自我改进 / 长上下文规划 / Agent 架构
2. CritICL — 利用小模型失败模式做推理时弱到强泛化
arXiv: https://arxiv.org/abs/2608.27455
核心洞察:同一模型家族中,小模型的失败模式在大模型上也呈现结构化规律,而不是随机噪声。CritICL 利用弱模型的失败模式作为引导信号,在推理时(inference-time)提升强模型推理能力,无需重复采样或依赖外部验证器。相比传统推理时缩放方法(BoN、PoT),更高效。
关联主题: 推理时缩放 / RAG 质量增强 / 模型协作
3. What Does an Evaluation License? — Inspect Evals 的 Claim-Replay 问题
arXiv: https://arxiv.org/abs/2608.19269
评测工件声明的指标未必能真正证明该声明——因为复现所需的历史证据和语义基础可能并不可用。本文对 124 个 Inspect Evals 单元做系统普查,发现其中 110 个在确定性推理前就停住了,因为所需的历史证据或语义基础不可用。评测本身的可复现性是 Agent/RAG 系统评估的重要前提。
关联主题: 评测可信度 / Benchmark 可靠性 / Agent 评估
📬 Substack 条目(1 条)
Agentic RAG vs CUA vs A2A: 哪个模式你需要?
来源: theaiengineer.substack.com · 2026
实用框架对比: - Agentic RAG:Orchestrator Agent 用 MCP 查询内部数据库,适合知识检索场景 - CUA(Computer Use Agent):处理遗留系统的表单填写、GUI 交互 - A2A:跨供应商 Agent 协调协议(但 2026 年底前生产级仍有风险)
2026 年末成熟企业系统可能是三者融合:Orchestrator 统筹、MCP 连接数据源、A2A 打通多供应商 Agent。RAG 不会死,它在向 Agentic RAG / GraphRAG / Hybrid Retrieval 演进。
关联主题: Agentic RAG / Agent 架构选型
📌 去重说明
本期候选均来自 2026-08-24~26 发布论文,与最近 7 天雷达文件无重复。Substack 条目为 theaiengineer 新帖,非历史已覆盖内容。
Tom 文献雷达 · 每日 3 次 · 轻量版 生成时间:2026-08-29 20:40 CST