SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

  • 类型:arxiv
  • 标识:2609.01343
  • 链接:https://arxiv.org/abs/2609.01343
  • 主分类:rag
  • 形态:method
  • TLDR:Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedd
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-02-agent-rag-longcontext-candidates.json