Lune

NeurIPS2025Top-tier venue

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Yuancheng Wang, Jiachen Zheng, Junan Zhang, Xueyao Zhang, Huan Liao, Zhizheng Wu

2025Year
25Citations
2Top-tier citations

Abstract

We introduce Metis, a foundation model for unified speech generation. Unlike previous task-specific or multi-task models, Metis follows a pre-training and fine-tuning paradigm. It is pre-trained on large-scale unlabeled speech data using masked generative modeling and then fine-tuned to adapt to diverse speech generation tasks. Specifically, 1) Metis utilizes two discrete speech representations: SSL tokens derived from speech self-supervised learning (SSL) features, and acoustic tokens directly quantized from waveforms. 2) Metis performs masked generative pretraining on SSL tokens, utilizing 300K hours of diverse speech data, without any additional condition. 3) Through fine-tuning with task-specific conditions, Metis achieves efficient adaptation to various speech generation tasks while supporting multimodal input, even when using limited data and trainable parameters. Experiments demonstrate that Metis can serve as a foundation model for unified speech generation: Metis outperforms state-of-the-art task-specific or multi-task systems across five speech generation tasks, including zero-shot text-to-speech, voice conversion, target speaker extraction, speech enhancement, and lip-to-speech, even with fewer than 20M trainable parameters or 300 times less training data. Audio samples are are available at https://metis-demo.github.io/. We release the code and model checkpoints at https://github.com/open-mmlab/Amphion.

Based on the above discussion, we propose Metis, a foundation model for unified speech generation. Specifically, Metis has the following key features: Masked Generative Pre-Training: Metis performs masked generative pre-training on SSL tokens using large-scale unlabeled speech data without any task-specific condition, as illustrated in Figure 1b. This pre-training phase establishes a strong foundation, allowing Metis to efficiently adapt to various downstream tasks through fine-tuning with minimal task-specific data.

Efficient Adaption to Various Speech Generation Tasks: Metis can be efficiently adapted to a wide range of speech generation tasks, including zero-shot text-to-speech, voice conversion, speech enhancement, target speaker extraction, and lip-to-speech, as illustrated in Figure 1c. The fine-tuned models achieve state-of-the-art results, even when using fewer than 20M trainable parameters or significantly less training data. Support for Multimodal Conditional Inputs: Metis supports multimodal conditional inputs during the fine-tuning phase, including text, audio, and video. This capability solidifies Metis as a foundation model for supporting various speech generation tasks. We also explore multi-task fine-tuning, showing that the pre-trained model can be efficiently adapted into a powerful multi-task model with minimal modification, enabling novel applications such as text-guided target speaker extraction with task combinations.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext dc3a6b37-241c-4956-bda0-22fd731f4e24

Cited by top-tier papers2

Ask how each one uses it

Builds on23

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines