Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
Yuancheng Wang, Jiachen Zheng, Junan Zhang, Xueyao Zhang, Huan Liao, Zhizheng Wu
Abstract
We introduce Metis, a foundation model for unified speech generation. Unlike previous task-specific or multi-task models, Metis follows a pre-training and fine-tuning paradigm. It is pre-trained on large-scale unlabeled speech data using masked generative modeling and then fine-tuned to adapt to diverse speech generation tasks. Specifically, 1) Metis utilizes two discrete speech representations: SSL tokens derived from speech self-supervised learning (SSL) features, and acoustic tokens directly quantized from waveforms. 2) Metis performs masked generative pretraining on SSL tokens, utilizing 300K hours of diverse speech data, without any additional condition. 3) Through fine-tuning with task-specific conditions, Metis achieves efficient adaptation to various speech generation tasks while supporting multimodal input, even when using limited data and trainable parameters. Experiments demonstrate that Metis can serve as a foundation model for unified speech generation: Metis outperforms state-of-the-art task-specific or multi-task systems across five speech generation tasks, including zero-shot text-to-speech, voice conversion, target speaker extraction, speech enhancement, and lip-to-speech, even with fewer than 20M trainable parameters or 300 times less training data. Audio samples are are available at https://metis-demo.github.io/. We release the code and model checkpoints at https://github.com/open-mmlab/Amphion.
Based on the above discussion, we propose Metis, a foundation model for unified speech generation. Specifically, Metis has the following key features: Masked Generative Pre-Training: Metis performs masked generative pre-training on SSL tokens using large-scale unlabeled speech data without any task-specific condition, as illustrated in Figure 1b. This pre-training phase establishes a strong foundation, allowing Metis to efficiently adapt to various downstream tasks through fine-tuning with minimal task-specific data.
Efficient Adaption to Various Speech Generation Tasks: Metis can be efficiently adapted to a wide range of speech generation tasks, including zero-shot text-to-speech, voice conversion, speech enhancement, target speaker extraction, and lip-to-speech, as illustrated in Figure 1c. The fine-tuned models achieve state-of-the-art results, even when using fewer than 20M trainable parameters or significantly less training data. Support for Multimodal Conditional Inputs: Metis supports multimodal conditional inputs during the fine-tuning phase, including text, audio, and video. This capability solidifies Metis as a foundation model for supporting various speech generation tasks. We also explore multi-task fine-tuning, showing that the pre-trained model can be efficiently adapted into a powerful multi-task model with minimal modification, enabling novel applications such as text-guided target speaker extraction with task combinations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc3a6b37-241c-4956-bda0-22fd731f4e24Cited by top-tier papers2
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language ModelingYuancheng Wang, Dekun Chen, Xueyao Zhang, Junan Zhang et al.NeurIPS 2025 · 22 citations
- Multi-Metric Preference Alignment for Generative Speech RestorationJunan Zhang, Xueyao Zhang, Jing Yang, Yuancheng Wang et al.AAAI 2026 · 6 citations
Builds on23
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
Related papers
- Generative Pre-training for Speech with Flow MatchingAlexander H. Liu, Matthew Le, Apoorv Vyas, Bowen Shi et al.ICLR 2024 · 66 citations
- UniAudio: Towards Universal Audio Generation with Large Language ModelsDongchao Yang, Jinchuan Tian, Xu Tan, Rongjie Huang et al.ICML 2024 · 54 citations
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec TransformerYuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng et al.ICLR 2025
- UniWav: Towards Unified Pre-training for Speech Representation Learning and GenerationAlexander H. Liu, Sang-gil Lee, Chao-Han Huck Yang, Yuan Gong et al.ICLR 2025
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang et al.ACL 2022 · 104 citations
