Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

  • 类型:arxiv
  • 标识:2608.03796
  • 链接:https://arxiv.org/abs/2608.03796
  • 主分类:engineering
  • 形态:application
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:A practitioner's study of how to make distillation training efficient is presented, organised around two systems contributions, and a fused, chunked KL loss is introduced, making peak memory linear in the sequence length.
  • 待LLM分类:否
  • 标题中文:高效 LLM 知识蒸馏:离线 Top-K Logits 与融合分块 KL 损失
  • TLDR中文:本文是一项关于如何提升蒸馏训练效率的实践研究,围绕两项系统贡献展开,并提出一种融合的 chunked KL loss,使峰值内存随序列长度线性增长。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-10-agent-rag-longcontext-candidates.json
  • [S2 enrich]