CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls
Li Chai, Donglin Wang
摘要
Lyric-to-melody generation is a highly challenging task in the field of AI music generation. Due to the difficulty of learning strict yet weak correlations between lyrics and melodies, previous methods have suffered from weak controllability, low-quality and poorly structured generation. To address these challenges, we propose CSL-L2M, a controllable song-level lyric-to-melody generation method based on an in-attention Transformer decoder with fine-grained lyric and musical controls, which is able to generate full-song melodies matched with the given lyrics and user-specified musical attributes. Specifically, we first introduce REMI-Aligned, a novel music representation that incorporates strict syllable- and sentence-level alignments between lyrics and melodies, facilitating precise alignment modeling. Subsequently, sentence-level semantic lyric embeddings independently extracted from a sentence-wise Transformer encoder are combined with word-level part-of-speech embeddings and syllable-level tone embeddings as fine-grained controls to enhance the controllability of lyrics over melody generation. Then we introduce human-labeled musical tags, sentence-level statistical musical attributes, and learned musical features extracted from a pre-trained VQ-VAE as coarse-grained, fine-grained and high-fidelity controls, respectively, to the generation process, thereby enabling user control over melody generation. Finally, an in-attention Transformer decoder technique is leveraged to exert fine-grained control over the full-song melody generation with the aforementioned lyric and musical conditions. Experimental results demonstrate that our proposed CSL-L2M outperforms the state-of-the-art models, generating melodies with higher quality, better controllability and enhanced structure.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsYu-Siang Huang, Yi-Hsuan YangACM MM 2020 · 被引用 265 次
- SongMASS: Automatic Song Writing with Pre-training and Alignment ConstraintZhonghao Sheng, Kaitao Song, Xu Tan, Yi Ren 等AAAI 2021 · 被引用 84 次
- TeleMelody: Lyric-to-Melody Generation with a Template-Based Two-Stage MethodZeqian Ju, Peiling Lu, Xu Tan, Rui Wang 等EMNLP 2022 · 被引用 19 次
- ReLyMe: Improving Lyric-to-Melody Generation by Incorporating Lyric-Melody RelationshipsChen Zhang, LuChin Chang, Songruoyao Wu, Xu Tan 等ACM MM 2022 · 被引用 13 次
相关 Paper
- SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-TrainingJiaxing Yu, Xinda Wu, Yunfei Xu, Tieyao Zhang 等AAAI 2025 · 被引用 2 次
- AI-Lyricist: Generating Music and Vocabulary Constrained LyricsXichu Ma, Ye Wang, Min-Yen Kan, Wee Sun LeeACM MM 2021 · 被引用 22 次
- LeVo: High-Quality Song Generation with Multi-Preference AlignmentShun Lei, Yaoxun Xu, Zhiwei Lin, Huaicheng Zhang 等NeurIPS 2025 · 被引用 43 次
- Unsupervised Melody-to-Lyrics GenerationYufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone 等ACL 2023 · 被引用 6 次
- MuseControlLite: Multifunctional Music Generation with Lightweight ConditionersFang-Duo Tsai, Shih-Lun Wu, Weijaw Lee, Sheng-Ping Yang 等ICML 2025
