Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

  • 类型:arxiv
  • 标识:2608.20492
  • 链接:https://arxiv.org/abs/2608.20492
  • 主分类:multimodal
  • 形态:method
  • TLDR:Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, post-training on large multi-task datasets remains challenging, as existing reinforcement learning methods sample on-policy groups with few high-quality rollouts even with costly chain-of-thought (CoT) generation. In this paper, we study the sample efficiency and scalability of RL post-training for video MLLMs and introduce OraRL. We identify an overlooked role for annotations: Beyond scoring rollouts, each can enter its on-policy group as an oracle rollout, a direct positive optim
  • 副分类:engineering
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-08-26-agent-rag-longcontext-candidates.json