Tom 文献雷达 · Agent RAG LongContext · 2026-08-31 08:40 UTC+8

本期候选(8条)

高价值 · Agent × World Model / Self-Improvement

Agentic Game Dev as a Verifiable Trajectory Data Engine agent | multimodal | systems - 核心论点:scaling world models 不能只靠更多视频 crawl,必须构建递归数据引擎提供 grounded reward signal。代码 agent 成功的原因:代码可执行,compiler/runtime 能给出高质量 RL reward。相比之下 spatial generation 仍依赖 CLIP 等模糊 proxy,难以支持 RL post-training。游戏开发提供了一种可行的 alternative。 - 关键信号:HF 178票,arXiv 2608.25518,2026-08-25

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents agent - 核心论点:现有 self-improvement 方法只在执行结束后处理 experience,无法 redirect 正在运行的 agent。提出 PILOT 框架:self-improvement 应该是 live 的,用 emerging experience 同时 redirect 当前运行 + 更新 persistent harness。 - 关键信号:HF 30票,arXiv 2608.26530,2026-08-26

中高价值 · RAG / Evaluation / Systems

Procedura: Agentic 3D Modeling with Procedural Control agent | rag | systems - 核心思路:3D shape as code,用 LLM coding 能力做 3D modeling。Agent 从 text prompt 规划 object 为 assembly graph,写成 parametric program。解决 dense mesh 无 part decomposition、无法编辑参数的问题。 - 关键信号:HF 11票,tag 同时含 rag,2026-08-25

CritICL: Inference-Time Weak-to-Strong Generalization from Small LM Failure Modes rag | systems - 核心思路:LLM failure modes 在同家族不同 scale 间有结构化 pattern。CritICL 利用弱模型的 failure modes 作为 guidance,在 inference time 改进推理,同时保持高效率(不依赖 repeated generation 或 external verifier)。 - 关键信号:HF 9票,tag 含 rag,2026-08-26

What Does an Evaluation License? — Commit-Bound Census of Inspect Evals benchmark | systems - 核心论点:评估 artifact 只规定 forward computation(task、scorer、metric),不等于 license 了那个 claim——因为 replay 所需的 historical evidence 和语义 grounding 可能 unbound。对 124 个 Inspect Evals 单元做机械审查:110 个止步于确定性推理之前。 - 关键信号:benchmark 视角稀缺,HF 3票,2026-08-24

一般候选(可扫)

标题 标签 信号
Luce: Relightable Gaussians for 3D Asset Generation multimodal HF 12票
TacForcing: Streaming Action + Tactile Feedback multimodal HF 5票
EditaLive! Unified Character Video Editing for Live Streaming multimodal HF 3票

Substack 线索

Agent Memory Is Not RAG: A Practical Map for Building Long-Horizon AI Agents — Claudio Stamile,2026-01

Agent memory overlaps with RAG、context engineering 和 "LLM memory",但 intent 不同:它 built to persist, evolve, and support long-horizon behavior。提出三个正交维度:forms、functions、lifecycles,取代简单的 short vs long-term 二分法。

RAG is about retrieval quality; long-context is about sequence handling; agent memory is about adaptation over time.

价值: 概念辨析清晰,适合用来检验 agent memory 设计思路。


Tom 文献雷达 · 每轮最多 8 条候选 · 高价值 3-4 条 · Substack ≤1 条 · 正文 ≤1200 字