FT-GAN: Fine-Grained Tune Modeling for Chinese Opera Synthesis
Meizhen Zheng, Peng Bai, Xiaodong Shi, Xun Zhou, Yiting Yan
Abstract
Although singing voice synthesis (SVS) has made significant progress recently, with its unique styles and various genres, Chinese opera synthesis requires greater attention but is rarely studied for lack of training data and high expressiveness. In this work, we build a high-quality Gezi Opera (a type of Chinese opera popular in Fujian and Taiwan) audio-text alignment dataset and formulate specific data annotation methods applicable to Chinese operas. We propose FT-GAN, an acoustic model for fine-grained tune modeling in Chinese opera synthesis based on the empirical analysis of the differences between Chinese operas and pop songs. To further improve the quality of the synthesized opera, we propose a speech pre-training strategy for additional knowledge injection. The experimental results show that FT-GAN outperforms the strong baselines in SVS on the Gezi Opera synthesis task. Extensive experiments further verify that FT-GAN performs well on synthesis tasks of other operas such as Peking Opera. Audio samples, the dataset, and the codes are available at https://zhengmidon.github.io/FTGAN.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df0fffdc-e08e-4132-8613-118e3e86461fBuilds on5
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing SynthesizersKai Shen, Zeqian Ju, Xu Tan, Eric Liu et al.ICLR 2024 · 362 citations
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen et al.AAAI 2022 · 348 citations
- CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency ModelZhen Ye, Wei Xue, Xu Tan, Jie Chen et al.ACM MM 2023 · 30 citations
- UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice SynthesisYi Lei, Shan Yang, Xinsheng Wang, Qicong Xie et al.AAAI 2023 · 15 citations
Related papers
- Text-to-Song: Towards Controllable Music Generation Incorporating Vocal and AccompanimentZhiqing Hong, Rongjie Huang, Xize Cheng, Yongqi Wang et al.ACL 2024 · 3 citations
- Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local ModelingPeng Bai, Yue Zhou, Meizhen Zheng, Wujin Sun et al.EMNLP 2023 · 4 citations
- Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale CorpusRongjie Huang, Feiyang Chen, Yi Ren, Jinglin Liu et al.ACM MM 2021 · 75 citations
- TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow MatchingWenxiang Guo, Yu Zhang, Changhao Pan, Rongjie Huang et al.AAAI 2025 · 21 citations
- DeepSinger: Singing Voice Synthesis with Data Mined From the WebYi Ren, Xu Tan, Tao Qin, Jian Luan et al.KDD 2020 · 72 citations
