视频选题榜 2026-07-21

  • https://arxiv.org/abs/2605.21384 · SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents · 把 reward hacking 做成 0/1 可测的 hacking gap · round=R1
  • https://arxiv.org/abs/2607.06624 · AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation · 把编码 Agent 70% 分数拆成可读评审 — 形式化验证做硬检查 + LLM 评审解释扣分 · round=R2
  • https://arxiv.org/abs/2607.16169 · When Does Muon Help Agentic Reinforcement Learning? · Muon 不是 AdamW 即插即用替代品 — 0.290→0.901 三件套协同 (优化器×estimator×lr) · round=R3