Autoregressive Speech Synthesis without Vector Quantization
Lingwei Meng, Long Zhou, Shujie Liu, Sanyuan Chen, Bing Han, Shujie Hu, Yanqing Liu, Jinyu Li, Sheng Zhao, Xixin Wu, Helen M. Meng, Furu Wei
摘要
We present MELLE, a novel continuousvalued token based language modeling approach for text-to-speech synthesis (TTS). MELLE autoregressively generates continuous mel-spectrogram frames directly from text condition, bypassing the need for vector quantization, which is typically designed for audio compression and sacrifices fidelity compared to continuous representations. Specifically, (i) instead of cross-entropy loss, we apply regression loss with a proposed spectrogram flux loss function to model the probability distribution of the continuous-valued tokens; (ii) we have incorporated variational inference into MELLE to facilitate sampling mechanisms, thereby enhancing the output diversity and model robustness. Experiments demonstrate that, compared to the two-stage codec language model VALL-E and its variants, the single-stage MELLE mitigates robustness issues by avoiding the inherent flaws of sampling vector-quantized codes, achieves superior performance across multiple metrics, and, most importantly, offers a more streamlined paradigm. The demos of our work are provided at https://aka.ms/melle . 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Multimodal Latent Language Modeling with Next-Token DiffusionYutao Sun, Hangbo Bao, Wenhui Wang, Zhiliang Peng 等ICML 2026 · 被引用 54 次
- Hyperspherical Latents Improve Continuous-Token Autoregressive GenerationGuolin Ke, Hui XueICLR 2026 · 被引用 19 次
- Continuous Audio Language ModelsSimon Rouard, Manu Orsini, Axel Roebel, Neil Zeghidour 等ICLR 2026 · 被引用 13 次
- MotionStreamer: Streaming Motion Generation via Diffusion-Based Autoregressive Model in Causal Latent SpaceLixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan 等ICCV 2025 · 被引用 11 次
- OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow MatchingNghia-Huynh Nguyen-Hieu, Ngoc Son Nguyen, Huynh Nguyen Dang, Thieu Vo 等ACL 2025 · 被引用 7 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- Voicebox: Text-Guided Multilingual Universal Speech Generation at ScaleMatthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer 等NeurIPS 2023 · 被引用 613 次
- YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneEdresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior 等ICML 2022 · 被引用 602 次
相关 Paper
- KALL-E: Autoregressive Speech Synthesis with Next-Distribution PredictionKangxiang Xia, Xinfa Zhu, Jixun Yao, Wenjie Tian 等AAAI 2026 · 被引用 3 次
- Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech SynthesisWeiwei Lin, Chenhang HeICLR 2025
- FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow MatchingHui Wang, Shujie Liu, Lingwei Meng, Jinyu Li 等ACM MM 2025 · 被引用 1 次
- Efficient Speech Language Modeling via Energy Distance in Continuous Latent SpaceZhengrui Ma, Yang Feng, Chenze Shao, Fandong Meng 等NeurIPS 2025 · 被引用 5 次
- UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and VocodingChenpeng Du, Yiwei Guo, Feiyu Shen, Zhijun Liu 等AAAI 2024 · 被引用 64 次
