MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Yuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng, Haotian Guo, Jiachen Zheng, Qiang Zhang, Xueyao Zhang, Shunsi Zhang, Zhizheng Wu
摘要
The recent large-scale text-to-speech (TTS) systems are usually grouped as autoregressive and non-autoregressive systems. The autoregressive systems implicitly model duration but exhibit certain deficiencies in robustness and lack of duration controllability. Non-autoregressive systems require explicit alignment information between text and speech during training and predict durations for linguistic units (e.g. phone), which may compromise their naturalness. In this paper, we introduce Masked Generative Codec Transformer (MaskGCT), a fully non-autoregressive TTS model that eliminates the need for explicit alignment information between text and speech supervision, as well as phone-level duration prediction. MaskGCT is a two-stage model: in the first stage, the model uses text to predict semantic tokens extracted from a speech self-supervised learning (SSL) model, and in the second stage, the model predicts acoustic tokens conditioned on these semantic tokens. MaskGCT follows the mask-and-predict learning paradigm. During training, MaskGCT learns to predict masked semantic or acoustic tokens based on given conditions and prompts. During inference, the model generates tokens of a specified length in a parallel manner. Experiments with 100K hours of in-thewild speech demonstrate that MaskGCT outperforms the current state-of-the-art zero-shot TTS systems in terms of quality, similarity, and intelligibility. Audio samples are available at https://maskgct.github.io/ . We release our code and model checkpoints at https://github.com/open-mmlab/Amphion/blob/ main/models/tts/maskgct . Recently, masked generative transformers, a class of generative models, have achieved significant results in the fields of image [11, 12, 13] , video [14, 15] , and audio [16, 17, 18] generation, demonstrating potential comparable to or superior to autoregressive models or diffusion models. These models employ a mask-and-predict training paradigm and utilize iterative parallel decoding during Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechSiyi Zhou, Yiquan Zhou, Yi He, Xun Zhou 等AAAI 2026 · 被引用 63 次
- FocalCodec: Low-Bitrate Speech Coding via Focal Modulation NetworksLuca Della Libera, Francesco Paissan, Cem Subakan, Mirco RavanelliNeurIPS 2025 · 被引用 35 次
- MoonCast: High-Quality Zero-Shot Podcast GenerationZeqian Ju, Dongchao Yang, Kai Shen, Yichong Leng 等NeurIPS 2025 · 被引用 32 次
- SpeechJudge: Towards Human-Level Judgment for Speech NaturalnessXueyao Zhang, Chaoren Wang, Huan Liao, Ziniu Li 等ICLR 2026 · 被引用 32 次
- Metis: A Foundation Speech Generation Model with Masked Generative Pre-trainingYuancheng Wang, Jiachen Zheng, Junan Zhang, Xueyao Zhang 等NeurIPS 2025 · 被引用 25 次
它引用的顶会 Paper21
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez 等NeurIPS 2023 · 被引用 843 次
相关 Paper
- T2V2: A Unified Non-Autoregressive Model for Speech Recognition and Synthesis via Multitask LearningNabarun Goswami, Hanqin Wang, Tatsuya HaradaICLR 2025
- Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech SynthesisYifan Yang, Shujie Liu, Jinyu Li, Yuxuan Hu 等ACM MM 2025 · 被引用 1 次
- MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-SpeechShengpeng Ji, Ziyue Jiang, Hanting Wang, Jialong Zuo 等ACL 2024 · 被引用 1 次
- DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific FactorsKeon Lee, Dong Won Kim, Jaehyeon Kim, Seungjun Chung 等ICLR 2025
- HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech SynthesisSang-Hoon Lee, Seung-Bin Kim, Ji-Hyun Lee, Eunwoo Song 等NeurIPS 2022 · 被引用 81 次
