From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

  • 类型:arxiv
  • 标识:2609.21788
  • 链接:https://arxiv.org/abs/2609.21788
  • 主分类:engineering
  • 形态:method
  • TLDR:A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks. Collecting additional full-task demonstrations for supervised fine-tuning (SFT) requires operators to repeat behaviors the policy already performs well. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches struggle to solve long-horizon tasks using only sparse rewards. We present PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at these bottlenecks w
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-22-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-09-22-agent-rag-longcontext-candidates.json