DeepSinger: Singing Voice Synthesis with Data Mined From the Web
Yi Ren, Xu Tan, Tao Qin, Jian Luan, Zhou Zhao, Tie-Yan Liu
摘要
In this paper 1 , we develop DeepSinger, a multi-lingual multi-singer singing voice synthesis (SVS) system, which is built from scratch using singing training data mined from music websites. The pipeline of DeepSinger consists of several steps, including data crawling, singing and accompaniment separation, lyrics-to-singing alignment, data filtration, and singing modeling. Specifically, we design a lyrics-to-singing alignment model to automatically extract the duration of each phoneme in lyrics starting from coarse-grained sentence level to fine-grained phoneme level, and further design a multi-lingual multi-singer singing model based on a feed-forward Transformer to directly generate linear-spectrograms from lyrics, and synthesize voices using Griffin-Lim. DeepSinger has several advantages over previous SVS systems: 1) to the best of our knowledge, it is the first SVS system that directly mines training data from music websites, 2) the lyrics-to-singing alignment model further avoids any human efforts for alignment labeling and greatly reduces labeling cost, 3) the singing model based on a feed-forward Transformer is simple and efficient, by removing the complicated acoustic feature modeling in parametric synthesis and leveraging a reference encoder to capture the timbre of a singer from noisy singing data, and 4) it can synthesize singing voices in multiple languages and multiple singers. We evaluate DeepSinger on our mined singing dataset that consists of about 92 hours data from 89 singers on three languages (Chinese, Cantonese and English). The results demonstrate that with the singing data purely mined from the Web, DeepSinger can synthesize high-quality singing voices in terms of both pitch accuracy and voice naturalness 2 . CCS CONCEPTS • Computing methodologies → Natural language processing; • Applied computing → Sound and music computing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen 等AAAI 2022 · 被引用 348 次
- PopMAG: Pop Music Accompaniment GenerationYi Ren, Jinzheng He, Xu Tan, Tao Qin 等ACM MM 2020 · 被引用 91 次
- Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale CorpusRongjie Huang, Feiyang Chen, Yi Ren, Jinglin Liu 等ACM MM 2021 · 被引用 75 次
- SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice GenerationRongjie Huang, Chenye Cui, Feiyang Chen, Yi Ren 等ACM MM 2022 · 被引用 46 次
- Learning the Beauty in Songs: Neural Singing Voice BeautifierJinglin Liu, Chengxi Li, Yi Ren, Zhiying Zhu 等ACL 2022 · 被引用 25 次
它引用的顶会 Paper1
相关 Paper
- ExpressiveSinger: Multilingual and Multi-Style Score-based Singing Voice Synthesis with Expressive Performance ControlShuqi Dai, Ming-Yu Liu, Rafael Valle, Siddharth GururaniACM MM 2024 · 被引用 6 次
- TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style ControlYu Zhang, Ziyue Jiang, Ruiqi Li, Changhao Pan 等EMNLP 2024 · 被引用 4 次
- TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow MatchingWenxiang Guo, Yu Zhang, Changhao Pan, Rongjie Huang 等AAAI 2025 · 被引用 21 次
- DeepRapper: Neural Rap Generation with Rhyme and Rhythm ModelingLanqing Xue, Kaitao Song, Duocai Wu, Xu Tan 等ACL 2021
- UniSinger: Unified End-to-End Singing Voice Synthesis With Cross-Modality Information MatchingZhiqing Hong, Chenye Cui, Rongjie Huang, Lichao Zhang 等ACM MM 2023 · 被引用 3 次
