Structuring MoE Expert Selection for Agentic Reinforcement Learning

  • 类型:arxiv
  • 标识:2610.07332
  • 链接:https://arxiv.org/abs/2610.07332
  • 主分类:agent
  • 形态:method
  • TLDR:Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing oper
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-08-agent-rag-longcontext-candidates.json