SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice Generation
Rongjie Huang, Chenye Cui, Feiyang Chen, Yi Ren, Jinglin Liu, Zhou Zhao, Baoxing Huai, Zhefeng Wang
摘要
Deep generative models have achieved significant progress in speech synthesis to date, while high-fidelity singing voice synthesis is still an open problem for its long continuous pronunciation, rich high-frequency parts, and strong expressiveness. Existing neural vocoders designed for text-to-speech cannot directly be applied to singing voice synthesis because they result in glitches and poor high-frequency reconstruction. In this work, we propose SingGAN, a generative adversarial network designed for high-fidelity singing voice synthesis. Specifically, 1) to alleviate the glitch problem in the generated samples, we propose source excitation with the adaptive feature learning filters to expand the receptive field patterns and stabilize long continuous signal generation; and 2) SingGAN introduces global and local discriminators at different scales to enrich low-frequency details and promote high-frequency reconstruction; and 3) To improve the training efficiency, SingGAN includes auxiliary spectrogram losses and sub-band feature matching penalty loss. To the best of our knowledge, SingGAN is the first work designed toward high-fidelity singing voice vocoding. Our evaluation of SingGAN demonstrates the state-of-the-art results with higherquality (MOS 4.05) samples. Also, SingGAN enables a sample speed of 50x faster than real-time on a single NVIDIA 2080Ti GPU. We further show that SingGAN generalizes well to the mel-spectrogram inversion of unseen singers, and the end-to-end singing voice synthesis system SingGAN-SVS enjoys a two-stage pipeline to transform the music scores into expressive singing voices. Audio samples are available at https://SingGAN.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-SpeechRongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu 等ACM MM 2022 · 被引用 182 次
- GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-SpeechRongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui 等NeurIPS 2022 · 被引用 99 次
- SongCreator: Lyrics-based Universal Song GenerationShun Lei, Yixuan Zhou, Boshi Tang, Max W. Y. Lam 等NeurIPS 2024 · 被引用 33 次
- Contrastive Learning with Positive-Negative Frame Mask for Music RepresentationDong Yao, Zhou Zhao, Shengyu Zhang, Jieming Zhu 等WWW 2022 · 被引用 26 次
- TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow MatchingWenxiang Guo, Yu Zhang, Changhao Pan, Rongjie Huang 等AAAI 2025 · 被引用 21 次
它引用的顶会 Paper15
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- Grad-TTS: A Diffusion Probabilistic Model for Text-to-SpeechVadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova 等ICML 2021 · 被引用 715 次
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 被引用 663 次
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
相关 Paper
- Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale CorpusRongjie Huang, Feiyang Chen, Yi Ren, Jinglin Liu 等ACM MM 2021 · 被引用 75 次
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen 等AAAI 2022 · 被引用 348 次
- CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational AutoencoderJianwei Cui, Yu Gu, Shihao Chen, Jie Zhang 等AAAI 2025 · 被引用 1 次
- UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice SynthesisYi Lei, Shan Yang, Xinsheng Wang, Qicong Xie 等AAAI 2023 · 被引用 15 次
- BigVGAN: A Universal Neural Vocoder with Large-Scale TrainingSang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro 等ICLR 2023 · 被引用 46 次
