IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
- 类型:arxiv
- 标识:2609.21346
- 链接:https://arxiv.org/abs/2609.21346
- 主分类:engineering
- 形态:position
- TLDR:Mixture-of-Experts (MoE) scales capacity, but existing designs cannot set three quantities independently. For a single token, participation is how many experts contribute knowledge to its output, execution is how many are actually computed (compute cost), and materialization is how many expert-sized parameter sets must be built and stored (memory cost). Sparse routing keeps execution and materialization low, but shrinks participation: for each token, only a few experts contribute. Dense output-mixing restores full participation, but its execution grows with the number of experts. Parameter-mer
- 待LLM分类:是
- 标题中文:IntBMoE:将块级条件化融入专家组合的全参与 Mixture-of-Experts
- TLDR中文:Mixture-of-Experts (MoE) 可扩展容量,但现有设计无法同时独立设定三个量:对于单个 token,参与度(participation)是贡献知识的专家数,执行(execution)是实际计算的专家数(计算开销),物化(materialization)是需构建与存储的专家规模参数集数量(内存开销)。稀疏路由保持 execution 与 materialization 较低,但缩小了 participation:每个 token 只有少数专家参与。稠密输出混合恢复完全 participation,但 execution 随专家数增长。参数合并
- 来源文件:
- /inbox/tom/_candidates/2026-09-21-agent-rag-longcontext-candidates.json