Tom 文献雷达 · Agent RAG LongContext · 2026-08-30 14:40

本批次说明: 候选与 08:40 批次完全一致(HuggingFace Daily 同一批次推送)。本次补充 Substack 新线索。

候选摘要(8 条,与 08:40 批次相同)

# 标题 来源 日期 信号
1 Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models HF Daily 2026-08-25 ⭐⭐⭐ 135票
2 PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents HF Daily 2026-08-26 ⭐⭐ 28票
3 Procedura: Agentic 3D Modeling with Procedural Control HF Daily 2026-08-25 ⭐ 10票
4 CritICL: Inference-Time Weak-to-Strong Generalization from Small LM Failure Modes HF Daily 2026-08-26 ⭐ 8票
5 Luce: Relightable Gaussians for 3D Asset Generation HF Daily 2026-08-24 7票
6 TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback HF Daily 2026-08-25 5票
7 EditaLive! Unified Character Video Editing for Live Streaming HF Daily 2026-08-26 3票
8 What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference HF Daily 2026-08-24 2票

高价值条目(4 条)

⭐⭐⭐ Agentic Game Dev → RL Data Engine(135票,2026-08-25)

链接: https://arxiv.org/abs/2608.25518

Scaling world models 的核心瓶颈不是数据量,而是 reward signal 质量。代码 agent 成功因编译器/runtimes 提供可执行 reward;3D/spatial 生成仍依赖 CLIP 等模糊 proxy,难以支持 RL post-training。用游戏开发构建可验证 RL data engine 是继 code agent 之后最有潜力的训练数据方向。Agent + RL + World Model 三领域交叉,重点跟进。

⭐⭐ PILOT in the Loop:实时自我改进(28票,2026-08-26)

链接: https://arxiv.org/abs/2608.26530

现有 self-improvement 方法在执行结束后才处理 experience,无法重定向当前 run。本文主张 live self-improvement:在执行过程中同时 redirect 当前 run 并更新 persistent harness。架构上需解决单 agent 自纠正(任务执行+轨迹评估同 context)与 subagent 委托的矛盾。对长期 agent 的 memory/反思机制有直接启发。

⭐⭐ CritICL:推理时弱到强泛化(8票,2026-08-26)

链接: https://arxiv.org/abs/2608.27455

核心洞察:同系列模型中弱模型的 failure modes 在强模型上也有结构化规律。CritICL 利用弱模型 failure modes 作 guidance source,兼顾效率与性能。检索失败模式的结构化分析可用于设计更强的 retrieval-augmented reasoning 系统。

⭐ Eval Benchmark 深水区(2票,2026-08-24)

链接: https://arxiv.org/abs/2608.19269

124 个 Inspect Evals 单元全量审查:110 个在确定性推理前就停止,因所需历史证据或语义 grounding 不可用。Eval artifact 报出的 metric 不必然 license 其 claim——对评测驱动的 RAG/长上下文系统有重要警示:别迷信 benchmark 数字。

Substack 补充(1 条,新增)

Agentic RAG vs CUA vs A2A:三种模式怎么选? 来源:The AI Engineer · https://theaiengineer.substack.com/p/agentic-rag-vs-cua-vs-a2a

2026 年末成熟企业系统形态示例:orchestrator agent 用 MCP 查询内部数据库(Agentic RAG)→ 委托 CUA sub-agent 填表 → 通过 A2A 与供应商 procurement agent 协调。注意 A2A 是互操作协议而非编排框架,仍需 LangGraph/ADK/CrewAI 处理内部逻辑;跨厂商 A2A 生产落地预计 2026 年底前不成熟。


本批次 08:40 与 14:40 候选完全一致;14:40 补充 Substack 新线索。 候选 JSON:/shared/research-kb/inbox/tom/_candidates/2026-08-30-agent-rag-longcontext-candidates.json