Tom 文献雷达 · Agent / RAG / 长上下文 · 2026-08-29 08:40

本期候选(8条)

# 来源 标题 标签 信号
1 HF Daily Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models agent, multimodal, systems ⭐119票
2 HF Daily PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents agent ⭐25票
3 HF Daily Procedura: Agentic 3D Modeling with Procedural Control agent, rag, systems ⭐7票
4 HF Daily CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes rag, systems ⭐4票
5 HF Daily EditaLive! Unified Character Video Editing for Live Streaming multimodal, systems 2票
6 HF Daily TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback multimodal 4票
7 HF Daily Luce: Relightable Gaussians for 3D Asset Generation multimodal 6票
8 HF Daily What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals benchmark, systems 2票

高价值条目(3条)

★ Agent 自我改进 / World Model 数据引擎

  • 标题:Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
  • arXiv:2608.25518
  • 核心观点: scaling world models 光靠爬视频不够,需要一个递归数据引擎提供 grounded reward signal。代码 agent 成功是因为代码可执行,RL 后训练能拿到高质量奖励信号;而空间生成仍依赖 CLIP 等模糊 proxy,难以支持 RL 后训练。游戏开发提供了一种可验证的轨迹数据生产路径。
  • 关联:Agent × RL post-training × World Model scaling

★ 实时自我改进(Live Self-Improvement)

  • 标题:PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
  • arXiv:2608.26530
  • 核心观点:现有 self-improvement 方法只在执行结束后处理 experience,无法重定向当前运行。PILOT 提出 live self-improvement:边执行边用 emerging experience 重定向当前任务,同时更新 persistent harness。
  • 关联:Agent × Long-Horizon × 自我改进机制

★ 评测可复现性 census

  • 标题:What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals
  • arXiv:2608.19269
  • 核心观点:评测 artifacts 声明指标≠授权该指标对应的 claim,因为 replay 所需的历史证据和语义 grounding 可能缺失。对 124 个 Inspect Evals 单元做机械审查,110 个在确定性推理前就停了(缺历史证据或语义 grounding)。
  • 关联:Benchmark × 可复现性 × 评测诚信

快速点评

  • Agent + RL 是本期最清晰的脉络:无论是 world model 数据引擎还是 live self-improvement,都在解决"奖励信号质量不足"这个核心瓶颈。游戏开发和代码执行之所以被反复引用,是因为它们天然提供可验证的奖励。
  • 长上下文 相关工作(尤其是 PILOT)继续深化"边跑边学"范式,与传统 batch 后处理模式形成对比。
  • 评测诚信(Inspect Evals census)值得关注:随着 benchmark 数量爆发,claim vs. metric 的 gap 正在被系统性审计。

Tom 文献雷达 · 每日3次(08:40 / 14:40 / 20:40 UTC+8) 数据源:arXiv metadata + Hugging Face Daily · 最多8条候选 · 高价值≤4条