Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local Modeling
Peng Bai, Yue Zhou, Meizhen Zheng, Wujin Sun, Xiaodong Shi
摘要
Singing Voice Synthesis (SVS) strives to synthesize pleasing vocals based on music scores and lyrics. The current acoustic models based on Transformer usually process the entire sequence globally and use a simple L1 loss. However, this approach overlooks the significance of local modeling within the sequence and the local optimization of the hard-to-synthesize parts in the predicted mel-spectrogram. Consequently, the synthesized audio exhibits local incongruities (e.g., local pronunciation jitter or local noise). To address this problem, we propose two methods to enhance local modeling in the acoustic model. First, we devise a nearest neighbor local attention, where each phoneme token focuses only on the adjacent phoneme tokens located before and after it. Second, we propose a phoneme-level local adaptive weights loss function that enables the model to focus more on the hard-to-synthesize parts of the mel-spectrogram. We have verified the universality of our methods on public Chinese pop song and Hokkien Gezi Opera datasets. Extensive experiments have demonstrated the effectiveness of our methods, resulting in significant improvements in both objective and subjective evaluations when compared to the strong baselines. Our code and demonstration samples are available at https://github.com/baipeng1/SVSELM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen 等AAAI 2022 · 被引用 348 次
- SG-Net: Syntax-Guided Machine Reading ComprehensionZhuosheng Zhang, Yuwei Wu, Junru Zhou, Sufeng Duan 等AAAI 2020 · 被引用 192 次
- Equalized Focal Loss for Dense Long-Tailed Object DetectionBo Li, Yongqiang Yao, Jingru Tan, Gang Zhang 等CVPR 2022 · 被引用 132 次
相关 Paper
- FT-GAN: Fine-Grained Tune Modeling for Chinese Opera SynthesisMeizhen Zheng, Peng Bai, Xiaodong Shi, Xun Zhou 等AAAI 2024 · 被引用 12 次
- Learning the Beauty in Songs: Neural Singing Voice BeautifierJinglin Liu, Chengxi Li, Yi Ren, Zhiying Zhu 等ACL 2022 · 被引用 25 次
- Robust Singing Voice Transcription Serves SynthesisRuiqi Li, Yu Zhang, Yongqi Wang, Zhiqing Hong 等ACL 2024 · 被引用 6 次
- SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice GenerationRongjie Huang, Chenye Cui, Feiyang Chen, Yi Ren 等ACM MM 2022 · 被引用 46 次
- UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice SynthesisYi Lei, Shan Yang, Xinsheng Wang, Qicong Xie 等AAAI 2023 · 被引用 15 次
