NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech
- 类型:arxiv
- 标识:2609.31892
- 链接:https://arxiv.org/abs/2609.31892
- 主分类:multimodal
- 形态:method
- TLDR:While modern text-to-speech (TTS) systems generate highly natural speech and support inline non-verbal vocalization (NVV) tags, accurate control over these events remains challenging. A key gap is the lack of established post-training methods for non-verbal control in continuous autoregressive flow-matching TTS. To this end, we present NVAlign, a direct-gradient post-training framework for NVV tag-following in this architecture. We first perform supervised fine-tuning (SFT) of TTS models and an NVV-aware automatic speech recognition (NV-ASR) model on NVV-annotated speech, then freeze the NV-AS
- 副分类:engineering
- 待LLM分类:否
- 标题中文:NVAlign:面向连续自回归流匹配文本到语音系统中非言语控制的直接梯度优化
- TLDR中文:现代文本到语音(TTS)系统虽能生成高度自然的语音并支持内联非言语发声(NVV)标签,但对这类事件的精确控制仍具挑战。核心缺口在于缺乏针对连续自回归流匹配 TTS 中非言语控制的后训练方法。为此,我们提出 NVAlign,一种在该架构下实现 NVV 标签跟随的直接梯度后训练框架。我们首先在带 NVV 标注的语音上对 TTS 模型与一个具备 NVV 感知能力的自动语音识别(NV-ASR)模型进行监督微调(SFT),随后冻结 NV-AS
- 来源文件:
- /inbox/tom/_candidates/2026-09-30-agent-rag-longcontext-candidates.json