视频选题榜 2026-07-25

  • https://arxiv.org/abs/2607.21051 · Sample-Efficient Learning from Agent Experience · ED 把 Agent 交互历史蒸馏进模型权重 · 保留 64.8% ICL 增益 · Direct SFT 仅 3.8% · 749 SWE + 6 文字冒险跨域验证 · 9.6× RL 样本效率 · weight-level persistence · round=R1
  • https://arxiv.org/abs/2607.20785 · Robostral Navigate · 单目 RGB + 8B VLM + image-space waypoint · R2R-CE 77.4% / +10.5pp 单目优势 / +5.3pp 多相机优势 / RxR-CE 75.1% · prefix-caching 22× token 压缩 · tree-attention mask · 跨 wheeled/legged/aerial 三类机器人无需重训 · sensor-to-action 输出空间切换层 · round=R2
  • https://arxiv.org/abs/2607.21557 · OpenForgeRL: Train Harness-native Agents in Any Environment · proxy 解耦 harness × 训练栈 · ClawEval pass³ 31.7 / pass@3 55.9 / QwenClawBench 33.7 · OSWorld-Verified 37.7 / Online-Mind2Web 63.0 / WebVoyager 72.3 · harness-native RL train-deploy bridge · round=R3