When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

  • 类型:arxiv
  • 标识:2609.19671
  • 链接:https://arxiv.org/abs/2609.19671
  • 主分类:engineering
  • 形态:method
  • TLDR:Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, trading reduced computation on easy instances for accuracy loss on hard instances. We formulate efficient reasoning as an instance-adaptive computation allocation problem and propose When2Think, a post-training framework for hybrid reasoning that dynamically allocates computation based on problem difficulty. Our method introduces
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-18-agent-rag-longcontext-candidates.json