Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
Yifan Yang, Shujie Liu, Jinyu Li, Yuxuan Hu, Haibin Wu, Hui Wang, Jianwei Yu, Lingwei Meng, Haiyang Sun, Yanqing Liu, Yan Lu, Kai Yu, Xie Chen
摘要
Recent zero-shot text-to-speech (TTS) systems face a common dilemma: autoregressive (AR) models suffer from slow generation and lack duration controllability, while non-autoregressive (NAR) models lack temporal modeling and typically require complex designs. In this paper, we introduce a novel pseudo-autoregressive (PAR) codec language modeling approach that unifies AR and NAR modeling. Combining explicit temporal modeling from AR with parallel generation from NAR, PAR generates dynamic-length spans at fixed time steps. Building on PAR, we propose PALLE, a two-stage TTS system that leverages PAR for initial generation followed by NAR refinement. In the first stage, PAR progressively generates speech tokens along the time dimension, with each step predicting all positions in parallel but only retaining the left-most span. In the second stage, low-confidence tokens are iteratively refined in parallel, leveraging the global contextual information. Experiments demonstrate that PALLE, trained on LibriTTS, outperforms state-of-the-art systems trained on large-scale data, including F5-TTS, E2-TTS, and MaskGCT, on the LibriSpeech test-clean set in terms of speech quality, speaker similarity, and intelligibility, while achieving up to ten times faster inference speed. Audio samples are available at https://microsoft.com/research/project/vall-e-x/palle.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language ModelingYuancheng Wang, Dekun Chen, Xueyao Zhang, Junan Zhang 等NeurIPS 2025 · 被引用 22 次
- Efficient Speech Language Modeling via Energy Distance in Continuous Latent SpaceZhengrui Ma, Yang Feng, Chenze Shao, Fandong Meng 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper28
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
相关 Paper
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec TransformerYuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng 等ICLR 2025
- P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech PromptingSungwon Kim, Kevin J. Shih, Rohan Badlani, João Felipe Santos 等NeurIPS 2023 · 被引用 75 次
- ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence ReorderingYakun Song, Zhuo Chen, Xiaofei Wang, Ziyang Ma 等AAAI 2025 · 被引用 75 次
- MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-SpeechShengpeng Ji, Ziyue Jiang, Hanting Wang, Jialong Zuo 等ACL 2024 · 被引用 1 次
- Autoregressive Speech Synthesis without Vector QuantizationLingwei Meng, Long Zhou, Shujie Liu, Sanyuan Chen 等ACL 2025 · 被引用 94 次
