Tom 文献雷达 · Agent / RAG / 长上下文 · 2026-08-29 08:40
本期候选(8条)
| # | 来源 | 标题 | 标签 | 信号 |
|---|---|---|---|---|
| 1 | HF Daily | Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models | agent, multimodal, systems | ⭐119票 |
| 2 | HF Daily | PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents | agent | ⭐25票 |
| 3 | HF Daily | Procedura: Agentic 3D Modeling with Procedural Control | agent, rag, systems | ⭐7票 |
| 4 | HF Daily | CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes | rag, systems | ⭐4票 |
| 5 | HF Daily | EditaLive! Unified Character Video Editing for Live Streaming | multimodal, systems | 2票 |
| 6 | HF Daily | TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback | multimodal | 4票 |
| 7 | HF Daily | Luce: Relightable Gaussians for 3D Asset Generation | multimodal | 6票 |
| 8 | HF Daily | What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals | benchmark, systems | 2票 |
高价值条目(3条)
★ Agent 自我改进 / World Model 数据引擎
- 标题:Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
- arXiv:2608.25518
- 核心观点: scaling world models 光靠爬视频不够,需要一个递归数据引擎提供 grounded reward signal。代码 agent 成功是因为代码可执行,RL 后训练能拿到高质量奖励信号;而空间生成仍依赖 CLIP 等模糊 proxy,难以支持 RL 后训练。游戏开发提供了一种可验证的轨迹数据生产路径。
- 关联:Agent × RL post-training × World Model scaling
★ 实时自我改进(Live Self-Improvement)
- 标题:PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
- arXiv:2608.26530
- 核心观点:现有 self-improvement 方法只在执行结束后处理 experience,无法重定向当前运行。PILOT 提出 live self-improvement:边执行边用 emerging experience 重定向当前任务,同时更新 persistent harness。
- 关联:Agent × Long-Horizon × 自我改进机制
★ 评测可复现性 census
- 标题:What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals
- arXiv:2608.19269
- 核心观点:评测 artifacts 声明指标≠授权该指标对应的 claim,因为 replay 所需的历史证据和语义 grounding 可能缺失。对 124 个 Inspect Evals 单元做机械审查,110 个在确定性推理前就停了(缺历史证据或语义 grounding)。
- 关联:Benchmark × 可复现性 × 评测诚信
快速点评
- Agent + RL 是本期最清晰的脉络:无论是 world model 数据引擎还是 live self-improvement,都在解决"奖励信号质量不足"这个核心瓶颈。游戏开发和代码执行之所以被反复引用,是因为它们天然提供可验证的奖励。
- 长上下文 相关工作(尤其是 PILOT)继续深化"边跑边学"范式,与传统 batch 后处理模式形成对比。
- 评测诚信(Inspect Evals census)值得关注:随着 benchmark 数量爆发,claim vs. metric 的 gap 正在被系统性审计。
Tom 文献雷达 · 每日3次(08:40 / 14:40 / 20:40 UTC+8) 数据源:arXiv metadata + Hugging Face Daily · 最多8条候选 · 高价值≤4条