ExpressiveSinger: Multilingual and Multi-Style Score-based Singing Voice Synthesis with Expressive Performance Control
Shuqi Dai, Ming-Yu Liu, Rafael Valle, Siddharth Gururani
Abstract
Singing Voice Synthesis (SVS) has significantly advanced with deep generative models, achieving high audio quality but still struggling with musicality, mainly due to the lack of performance control over timing, dynamics, and pitch, which are essential for music expression. Additionally, integrating data and supporting diverse languages and styles in SVS remain challenging. To tackle these issues, this paper presents ExpressiveSinger, an SVS framework that leverages a cascade of diffusion models to generate realistic singing across multiple languages, styles, and techniques from scores and lyrics. Our approach begins with consolidating, cleaning, annotating, and processing public singing datasets, developing a multilingual phoneme set, and incorporating different musical styles and techniques. We then design methods for generating expressive performance control signals including phoneme timing, F0 curves, and amplitude envelopes, which enhance musicality and model consistency, introduce more controllability, and reduce data requirements. Finally, we generate mel-spectrograms and audio from performance control signals with style guidance and singer timbre embedding. Our models also enable trained singers to sing in new languages and styles. Several listening tests reveal both musicality and controllability of our generated singing compared with existing works and human singing. We release the data for future research. Demo: https://shuqid.net/expressive-singing-synthesis.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e65325ad-d455-4fdf-8f82-85307d523903Cited by top-tier papers1
Ask how each one uses itRelated papers
- TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow MatchingWenxiang Guo, Yu Zhang, Changhao Pan, Rongjie Huang et al.AAAI 2025 · 21 citations
- TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style ControlYu Zhang, Ziyue Jiang, Ruiqi Li, Changhao Pan et al.EMNLP 2024 · 4 citations
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen et al.AAAI 2022 · 348 citations
- UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice SynthesisYi Lei, Shan Yang, Xinsheng Wang, Qicong Xie et al.AAAI 2023 · 15 citations
- DeepSinger: Singing Voice Synthesis with Data Mined From the WebYi Ren, Xu Tan, Tao Qin, Jian Luan et al.KDD 2020 · 72 citations
