Robust Singing Voice Transcription Serves Synthesis
Ruiqi Li, Yu Zhang, Yongqi Wang, Zhiqing Hong, Rongjie Huang, Zhou Zhao
Abstract
Note-level Automatic Singing Voice Transcription (AST) converts singing recordings into note sequences, facilitating the automatic annotation of singing datasets for Singing Voice Synthesis (SVS) applications. Current AST methods, however, struggle with accuracy and robustness when used for practical annotation. This paper presents ROSVOT, the first robust AST model that serves SVS, incorporating a multi-scale framework that effectively captures coarse-grained note information and ensures fine-grained frame-level segmentation, coupled with an attention-based pitch decoder for reliable pitch prediction. We also established a comprehensive annotation-and-training pipeline for SVS to test the model in realworld settings. Experimental findings reveal that ROSVOT achieves state-of-the-art transcription accuracy with either clean or noisy inputs. Moreover, when trained on enlarged, automatically annotated datasets, the SVS model outperforms its baseline, affirming the capability for practical application. Audio samples are available at https://rosvot.github.io . Codes can be found at https://github.com/RickyL-2000/ROSVOT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e5ab962-c9e1-4bb2-ad4d-144910eb7c13Cited by top-tier papers3
- TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style ControlYu Zhang, Ziyue Jiang, Ruiqi Li, Changhao Pan et al.EMNLP 2024 · 4 citations
- ISDrama: Immersive Spatial Drama Generation through Multimodal PromptingYu Zhang, Wenxiang Guo, Changhao Pan, Zhiyuan Zhu et al.ACM MM 2025 · 1 citation
- Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion TransformerKe Lei, Yu Zhang, Changhao Pan, Xueyi Pu et al.ICML 2026 · 1 citation
Builds on9
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing SynthesizersKai Shen, Zeqian Ju, Xu Tan, Eric Liu et al.ICLR 2024 · 362 citations
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen et al.AAAI 2022 · 348 citations
Related papers
- TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow MatchingWenxiang Guo, Yu Zhang, Changhao Pan, Rongjie Huang et al.AAAI 2025 · 21 citations
- Elucidate Gender Fairness in Singing Voice TranscriptionXiangming Gu, Wei Zeng, Ye WangACM MM 2023 · 4 citations
- SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice GenerationRongjie Huang, Chenye Cui, Feiyang Chen, Yi Ren et al.ACM MM 2022 · 46 citations
- Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local ModelingPeng Bai, Yue Zhou, Meizhen Zheng, Wujin Sun et al.EMNLP 2023 · 4 citations
- ReconVAT: A Semi-Supervised Automatic Music Transcription Framework for Low-Resource Real-World DataKin Wai Cheuk, Dorien Herremans, Li SuACM MM 2021 · 28 citations
