Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery
- 类型:arxiv
- 标识:2609.27980
- 链接:https://arxiv.org/abs/2609.27980
- 主分类:llm-infra
- 形态:method
- TLDR:Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the {\tt whisper-large-v3-turbo} variant reduced the decoder from 32 to 4 layers, while Distill-Whisper similarly reduced the decoder to only 2 layers. Although some attention has been put towards reducing the size of the encoder, no approach has seen wide adoption. This could be due to the need for custom inference implementations to take advantage of the compressed model. We present an approach that
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-24-agent-rag-longcontext-candidates.json