SongMASS: Automatic Song Writing with Pre-training and Alignment Constraint
Zhonghao Sheng, Kaitao Song, Xu Tan, Yi Ren, Wei Ye, Shikun Zhang, Tao Qin
摘要
Automatic song writing aims to compose a song (lyric and/or melody) by machine, which is an interesting topic in both academia and industry. In automatic song writing, lyric-tomelody generation and melody-to-lyric generation are two important tasks, both of which usually suffer from the following challenges: 1) the paired lyric and melody data are limited, which affects the generation quality of the two tasks, considering a lot of paired training data are needed due to the weak correlation between lyric and melody; 2) Strict alignments are required between lyric and melody, which relies on specific alignment modeling. In this paper, we propose SongMASS to address the above challenges, which leverages masked sequence to sequence (MASS) pre-training and attention based alignment modeling for lyric-to-melody and melody-to-lyric generation. Specifically, 1) we extend the original sentence-level MASS pre-training to song level to better capture long contextual information in music, and use a separate encoder and decoder for each modality (lyric or melody); 2) we leverage sentence-level attention mask and token-level attention constraint during training to enhance the alignment between lyric and melody. During inference, we use a dynamic programming strategy to obtain the alignment between each word/syllable in lyric and note in melody. We pre-train SongMASS on unpaired lyric and melody datasets, and both objective and subjective evaluations demonstrate that SongMASS generates lyric and melody with significantly better quality than the baseline method without pretraining or alignment constraint.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Do Language Models Plagiarize?Jooyoung Lee, Thai Le, Jinghui Chen, Dongwon LeeWWW 2023 · 被引用 109 次
- Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationBotao Yu, Peiling Lu, Rui Wang, Wei Hu 等NeurIPS 2022 · 被引用 104 次
- ReLyMe: Improving Lyric-to-Melody Generation by Incorporating Lyric-Melody RelationshipsChen Zhang, LuChin Chang, Songruoyao Wu, Xu Tan 等ACM MM 2022 · 被引用 13 次
- PoetryDiffusion: Towards Joint Semantic and Metrical Manipulation in Poetry GenerationZhiyuan Hu, Chumin Liu, Yue Feng, Anh Tuan Luu 等AAAI 2024 · 被引用 11 次
- Fine-Grained Position Helps Memorizing More, a Novel Music Compound Transformer Model with Feature Interaction FusionZuchao Li, Ruhan Gong, Yineng Chen, Kehua SuAAAI 2023 · 被引用 11 次
它引用的顶会 Paper2
相关 Paper
- SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-TrainingJiaxing Yu, Xinda Wu, Yunfei Xu, Tieyao Zhang 等AAAI 2025 · 被引用 2 次
- Unsupervised Melody-to-Lyrics GenerationYufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone 等ACL 2023 · 被引用 6 次
- SongComposer: A Large Language Model for Lyric and Melody Generation in Song CompositionShuangrui Ding, Zihan Liu, Xiaoyi Dong, Pan Zhang 等ACL 2025
- CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical ControlsLi Chai, Donglin WangAAAI 2025 · 被引用 1 次
- S²MILE: Semantic-and-Structure-Aware Music-Driven Lyric GenerationMu You, Fang Zhang, Shuai Zhang, Linli XuAAAI 2025
