Robust Singing Voice Transcription Serves Synthesis
Ruiqi Li, Yu Zhang, Yongqi Wang, Zhiqing Hong, Rongjie Huang, Zhou Zhao
摘要
Note-level Automatic Singing Voice Transcription (AST) converts singing recordings into note sequences, facilitating the automatic annotation of singing datasets for Singing Voice Synthesis (SVS) applications. Current AST methods, however, struggle with accuracy and robustness when used for practical annotation. This paper presents ROSVOT, the first robust AST model that serves SVS, incorporating a multi-scale framework that effectively captures coarse-grained note information and ensures fine-grained frame-level segmentation, coupled with an attention-based pitch decoder for reliable pitch prediction. We also established a comprehensive annotation-and-training pipeline for SVS to test the model in realworld settings. Experimental findings reveal that ROSVOT achieves state-of-the-art transcription accuracy with either clean or noisy inputs. Moreover, when trained on enlarged, automatically annotated datasets, the SVS model outperforms its baseline, affirming the capability for practical application. Audio samples are available at https://rosvot.github.io . Codes can be found at https://github.com/RickyL-2000/ROSVOT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style ControlYu Zhang, Ziyue Jiang, Ruiqi Li, Changhao Pan 等EMNLP 2024 · 被引用 4 次
- ISDrama: Immersive Spatial Drama Generation through Multimodal PromptingYu Zhang, Wenxiang Guo, Changhao Pan, Zhiyuan Zhu 等ACM MM 2025 · 被引用 1 次
- Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion TransformerKe Lei, Yu Zhang, Changhao Pan, Xueyi Pu 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper9
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing SynthesizersKai Shen, Zeqian Ju, Xu Tan, Eric Liu 等ICLR 2024 · 被引用 362 次
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen 等AAAI 2022 · 被引用 348 次
相关 Paper
- TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow MatchingWenxiang Guo, Yu Zhang, Changhao Pan, Rongjie Huang 等AAAI 2025 · 被引用 21 次
- Elucidate Gender Fairness in Singing Voice TranscriptionXiangming Gu, Wei Zeng, Ye WangACM MM 2023 · 被引用 4 次
- SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice GenerationRongjie Huang, Chenye Cui, Feiyang Chen, Yi Ren 等ACM MM 2022 · 被引用 46 次
- Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local ModelingPeng Bai, Yue Zhou, Meizhen Zheng, Wujin Sun 等EMNLP 2023 · 被引用 4 次
- ReconVAT: A Semi-Supervised Automatic Music Transcription Framework for Low-Resource Real-World DataKin Wai Cheuk, Dorien Herremans, Li SuACM MM 2021 · 被引用 28 次
