Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation
Dongchan Min, Dong Bok Lee, Eunho Yang, Sung Ju Hwang
Abstract
With rapid progress in neural text-to-speech (TTS) models, personalized speech generation is now in high demand for many applications. For practical applicability, a TTS model should generate high-quality speech with only a few audio samples from the given speaker, that are also short in length. However, existing methods either require to fine-tune the model or achieve low adaptation quality without fine-tuning. In this work, we propose StyleSpeech, a new TTS model which not only synthesizes high-quality speech but also effectively adapts to new speakers. Specifically, we propose Style-Adaptive Layer Normalization (SALN) which aligns gain and bias of the text input according to the style extracted from a reference speech audio. With SALN, our model effectively synthesizes speech in the style of the target speaker even from a single speech audio. Furthermore, to enhance StyleSpeech's adaptation to speech from new speakers, we extend it to Meta-StyleSpeech by introducing two discriminators trained with style prototypes, and performing episodic training. The experimental results show that our models generate high-quality speech which accurately follows the speaker's voice with single short-duration (1-3 sec) speech audio, significantly outperforming baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-SpeechRongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu et al.ACM MM 2022 · 182 citations
- GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-SpeechRongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui et al.NeurIPS 2022 · 99 citations
- HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech SynthesisSang-Hoon Lee, Seung-Bin Kim, Ji-Hyun Lee, Eunwoo Song et al.NeurIPS 2022 · 81 citations
- Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisZiyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He et al.ICLR 2024 · 75 citations
- P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech PromptingSungwon Kim, Kevin J. Shih, Rohan Badlani, João Felipe Santos et al.NeurIPS 2023 · 75 citations
Builds on2
Related papers
- AdaSpeech: Adaptive Text to Speech for Custom VoiceMingjian Chen, Xu Tan, Bohan Li, Yanqing Liu et al.ICLR 2021 · 79 citations
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language ModelsYinghao Aaron Li, Cong Han, Vinay S. Raghavan, Gavin Mischler et al.NeurIPS 2023 · 324 citations
- ArtSpeech: Adaptive Text-to-Speech Synthesis with Articulatory RepresentationsZhongxu Wang, Yujia Wang, Mingzhu Li, Hua HuangACM MM 2024 · 2 citations
- Multi-SpectroGAN: High-Diversity and High-Fidelity Spectrogram Generation with Adversarial Style Combination for Speech SynthesisSang-Hoon Lee, Hyun-Wook Yoon, Hyeong-Rae Noh, Ji-Hoon Kim et al.AAAI 2021 · 60 citations
- ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style ControlShengpeng Ji, Qian Chen, Wen Wang, Jialong Zuo et al.ACL 2025
