arXiv:2609.38658 · 多模态
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Tacit-TTS:从自回归解码到掩码预测的高效免转录语音克隆
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
- 类型:arxiv
- 标识:2609.38658
- 链接:https://arxiv.org/abs/2609.38658
- 主分类:multimodal
- 形态:method
- TLDR:TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generati
- 待LLM分类:否
- 标题中文:Tacit-TTS:从自回归解码到掩码预测的高效免转录语音克隆
- TLDR中文:基于自回归语义建模的 TTS 系统已展现出强大的零样本语音克隆能力与丰富的表现力变化,但其顺序解码带来显著延迟。非自回归替代方案生成速度更快,却通常依赖更具限制性的参考条件,例如推理时需要参考语音的转录文本。本文提出 Tacit-TTS,一种从 IndexTTS2 蒸馏而来的高效免转录零样本语音克隆系统。该模型将自回归 text-to-semantic 解码替换为掩码非自回归生成……
- 来源文件:
- /inbox/tom/_candidates/2026-10-02-agent-rag-longcontext-candidates.json