Tom 文献雷达 · Agent RAG LongContext · 2026-08-30 14:40
本批次说明: 候选与 08:40 批次完全一致(HuggingFace Daily 同一批次推送)。本次补充 Substack 新线索。
候选摘要(8 条,与 08:40 批次相同)
| # | 标题 | 来源 | 日期 | 信号 |
|---|---|---|---|---|
| 1 | Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models | HF Daily | 2026-08-25 | ⭐⭐⭐ 135票 |
| 2 | PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents | HF Daily | 2026-08-26 | ⭐⭐ 28票 |
| 3 | Procedura: Agentic 3D Modeling with Procedural Control | HF Daily | 2026-08-25 | ⭐ 10票 |
| 4 | CritICL: Inference-Time Weak-to-Strong Generalization from Small LM Failure Modes | HF Daily | 2026-08-26 | ⭐ 8票 |
| 5 | Luce: Relightable Gaussians for 3D Asset Generation | HF Daily | 2026-08-24 | 7票 |
| 6 | TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback | HF Daily | 2026-08-25 | 5票 |
| 7 | EditaLive! Unified Character Video Editing for Live Streaming | HF Daily | 2026-08-26 | 3票 |
| 8 | What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference | HF Daily | 2026-08-24 | 2票 |
高价值条目(4 条)
⭐⭐⭐ Agentic Game Dev → RL Data Engine(135票,2026-08-25)
链接: https://arxiv.org/abs/2608.25518
Scaling world models 的核心瓶颈不是数据量,而是 reward signal 质量。代码 agent 成功因编译器/runtimes 提供可执行 reward;3D/spatial 生成仍依赖 CLIP 等模糊 proxy,难以支持 RL post-training。用游戏开发构建可验证 RL data engine 是继 code agent 之后最有潜力的训练数据方向。Agent + RL + World Model 三领域交叉,重点跟进。
⭐⭐ PILOT in the Loop:实时自我改进(28票,2026-08-26)
链接: https://arxiv.org/abs/2608.26530
现有 self-improvement 方法在执行结束后才处理 experience,无法重定向当前 run。本文主张 live self-improvement:在执行过程中同时 redirect 当前 run 并更新 persistent harness。架构上需解决单 agent 自纠正(任务执行+轨迹评估同 context)与 subagent 委托的矛盾。对长期 agent 的 memory/反思机制有直接启发。
⭐⭐ CritICL:推理时弱到强泛化(8票,2026-08-26)
链接: https://arxiv.org/abs/2608.27455
核心洞察:同系列模型中弱模型的 failure modes 在强模型上也有结构化规律。CritICL 利用弱模型 failure modes 作 guidance source,兼顾效率与性能。检索失败模式的结构化分析可用于设计更强的 retrieval-augmented reasoning 系统。
⭐ Eval Benchmark 深水区(2票,2026-08-24)
链接: https://arxiv.org/abs/2608.19269
124 个 Inspect Evals 单元全量审查:110 个在确定性推理前就停止,因所需历史证据或语义 grounding 不可用。Eval artifact 报出的 metric 不必然 license 其 claim——对评测驱动的 RAG/长上下文系统有重要警示:别迷信 benchmark 数字。
Substack 补充(1 条,新增)
Agentic RAG vs CUA vs A2A:三种模式怎么选? 来源:The AI Engineer · https://theaiengineer.substack.com/p/agentic-rag-vs-cua-vs-a2a
2026 年末成熟企业系统形态示例:orchestrator agent 用 MCP 查询内部数据库(Agentic RAG)→ 委托 CUA sub-agent 填表 → 通过 A2A 与供应商 procurement agent 协调。注意 A2A 是互操作协议而非编排框架,仍需 LangGraph/ADK/CrewAI 处理内部逻辑;跨厂商 A2A 生产落地预计 2026 年底前不成熟。
本批次 08:40 与 14:40 候选完全一致;14:40 补充 Substack 新线索。
候选 JSON:/shared/research-kb/inbox/tom/_candidates/2026-08-30-agent-rag-longcontext-candidates.json