Drop the Beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
Ziqian Ning, Shuai Wang, Yuepeng Jiang, Jixun Yao, Lei He, Shifeng Pan, Jie Ding, Lei Xie
Abstract
Rap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits flexibility. In contrast, rap typically features simpler melodies, with a core focus on a strong rhythmic sense that harmonizes with accompanying beats. In this paper, we propose Freestyler, the first system that generates rapping vocals directly from lyrics and accompaniment inputs. Freestyler utilizes language model-based token generation, followed by a conditional flow matching model to produce spectrograms and a neural vocoder to restore audio. It allows a 3-second prompt to enable zero-shot timbre control. Due to the scarcity of publicly available rap datasets, we also present RapBank, a rap song dataset collected from the internet, alongside a meticulously designed processing pipeline. Experimental results show that Freestyler produces high-quality rapping voice generation with enhanced naturalness and strong alignment with accompanying beats, both stylistically and rhythmically.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3cf9bcf-743d-4305-a756-4cedf0b5514bCited by top-tier papers2
- UniVocal: Unified Speech-Singing Code-Switching SynthesisYufei Shi, Qian Chen, Wen Wang, Xiangang Li et al.ACL 2026
- SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song GenerationZihan Liu, Shuangrui Ding, Zhixiong Zhang, Xiaoyi Dong et al.ICML 2025
Builds on11
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar et al.NeurIPS 2023 · 910 citations
- Grad-TTS: A Diffusion Probabilistic Model for Text-to-SpeechVadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova et al.ICML 2021 · 715 citations
- NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing SynthesizersKai Shen, Zeqian Ju, Xu Tan, Eric Liu et al.ICLR 2024 · 362 citations
Related papers
- DeepRapper: Neural Rap Generation with Rhyme and Rhythm ModelingLanqing Xue, Kaitao Song, Duocai Wu, Xu Tan et al.ACL 2021
- RapVerse: Coherent Vocals and Whole-Body Motion Generation from TextJiaben Chen, Xin Yan, Yihang Chen, Siyuan Cen et al.ICCV 2025 · 7 citations
- SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task EditorChenyu Yang, Shuai Wang, Hangting Chen, Jianwei Yu et al.AAAI 2025 · 9 citations
- SegTune: Structured and Fine-Grained Control for Song GenerationYuejiao Wang, Zihao Ji, Pengfei Cai, Xu Li et al.ACL 2026 · 2 citations
- LeVo: High-Quality Song Generation with Multi-Preference AlignmentShun Lei, Yaoxun Xu, Zhiwei Lin, Huaicheng Zhang et al.NeurIPS 2025 · 43 citations
