X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

  • 类型:arxiv
  • 标识:2609.11412
  • 链接:https://arxiv.org/abs/2609.11412
  • 主分类:multimodal
  • 形态:method
  • TLDR:Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during
  • 副分类:rag
  • 待LLM分类:否
  • 标题中文:X-AuT:面向语音 LLM 的渐进式音频编码器压缩与跨尺度蒸馏
  • TLDR中文:缩减音频编码器深度可降低语音大语言模型的推理成本,但移除完整块会扰动解码器所消费的 embedding,并可能导致删除和过早序列结束错误。我们提出 X-AuT,一个通过短行为探针选择层组合,并通过表征对齐、跨尺度蒸馏、学生策略监督调度以及 LoRA 微调来恢复剪枝模型的渐进式框架。语言模型主干保持冻结,而 attention LoRA adapter 与绑定的输出 embedding 在训练期间自适应。
  • 来源文件
  • /inbox/tom/_candidates/2026-09-11-agent-rag-longcontext-candidates.json