Switch Transformers: Scaling to Trillion Parameter Models with Simple\n and Efficient Sparsity

  • 类型:arxiv
  • 标识:2101.03961
  • 链接:https://arxiv.org/abs/2101.03961
  • 主题:rag
  • 主分类:llm-infra
  • 形态:method
  • 被引:4614
  • 被引来源:Semantic Scholar
  • S2被引:4614
  • OpenAlex被引:710
  • 影响力被引:565
  • TLDR:This work simplifies the MoE routing algorithm and design intuitive improved models with reduced communication and computational costs and shows large sparse models may be trained, for the first time, with lower precision formats.
  • OpenAlex ID:W4287391717
  • OpenAlex DOI:10.48550/arxiv.2101.03961
  • DOI:10.48550/arxiv.2101.03961
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://arxiv.org/pdf/2101.03961
  • OpenAlex更新:2026-08-22
  • 待LLM分类:否
  • 成熟度:research
  • 场景:稀疏专家模型、大规模训练
  • 标题中文:Switch Transformers:通过简单且高效的稀疏性将模型扩展到万亿参数规模
  • TLDR中文:简化了 MoE 路由算法,设计出通信与计算成本更低的直观改进模型,并首次证明大型稀疏模型可以使用更低精度格式进行训练
  • 来源文件
  • [OpenAlex discover]
  • [S2 enrich]