X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
- 类型:arxiv
- 标识:2609.11412
- 链接:https://arxiv.org/abs/2609.11412
- 主分类:multimodal
- 形态:method
- TLDR:Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during
- 副分类:rag
- 待LLM分类:否
- 标题中文:X-AuT:面向语音 LLM 的渐进式音频编码器压缩与跨尺度蒸馏
- TLDR中文:缩减音频编码器深度可降低语音大语言模型的推理成本,但移除完整块会扰动解码器所消费的 embedding,并可能导致删除和过早序列结束错误。我们提出 X-AuT,一个通过短行为探针选择层组合,并通过表征对齐、跨尺度蒸馏、学生策略监督调度以及 LoRA 微调来恢复剪枝模型的渐进式框架。语言模型主干保持冻结,而 attention LoRA adapter 与绑定的输出 embedding 在训练期间自适应。
- 来源文件:
- /inbox/tom/_candidates/2026-09-11-agent-rag-longcontext-candidates.json