PopMAG: Pop Music Accompaniment Generation
Yi Ren, Jinzheng He, Xu Tan, Tao Qin, Zhou Zhao, Tie-Yan Liu
Abstract
In pop music, accompaniments are usually played by multiple instruments (tracks) such as drum, bass, string and guitar, and can make a song more expressive and contagious by arranging together with its melody. Previous works usually generate multiple tracks separately and the music notes from different tracks not explicitly depend on each other, which hurts the harmony modeling. To improve harmony, in this paper 1 , we propose a novel MUlti-track MIDI representation (MuMIDI), which enables simultaneous multi-track generation in a single sequence and explicitly models the dependency of the notes from different tracks. While this greatly improves harmony, unfortunately, it enlarges the sequence length and brings the new challenge of long-term music modeling. We further introduce two new techniques to address this challenge: 1) We model multiple note attributes (e.g., pitch, duration, velocity) of a musical note in one step instead of multiple steps, which can shorten the length of a MuMIDI sequence. 2) We introduce extra long-context as memory to capture long-term dependency in music. We call our system for pop music accompaniment generation as PopMAG. We evaluate PopMAG on multiple datasets (LMD, FreeMidi and CPMD, a private dataset of Chinese pop songs) with both subjective and objective metrics. The results demonstrate the effectiveness of PopMAG for multi-track harmony modeling and long-term context modeling. Specifically, PopMAG wins 42%/38%/40% votes when comparing with ground truth musical pieces on LMD, FreeMidi and CPMD datasets respectively and largely outperforms other state-ofthe-art music accompaniment generation models and multi-track MIDI representations in terms of subjective and objective metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5e84b18-b3be-44aa-9933-13fff61cd91bCited by top-tier papers17
- Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationBotao Yu, Peiling Lu, Rui Wang, Wei Hu et al.NeurIPS 2022 · 104 citations
- Video Background Music Generation with Controllable Music TransformerShangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang et al.ACM MM 2021 · 87 citations
- SongMASS: Automatic Song Writing with Pre-training and Alignment ConstraintZhonghao Sheng, Kaitao Song, Xu Tan, Yi Ren et al.AAAI 2021 · 84 citations
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao et al.ICCV 2023 · 51 citations
- Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion ModelsZiyu Wang, Lejun Min, Gus XiaICLR 2024 · 32 citations
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsYu-Siang Huang, Yi-Hsuan YangACM MM 2020 · 265 citations
- Encoding Musical Style with Transformer AutoencodersKristy Choi, Curtis Hawthorne, Ian Simon, Monica Dinculescu et al.ICML 2020 · 102 citations
- DeepSinger: Singing Voice Synthesis with Data Mined From the WebYi Ren, Xu Tan, Tao Qin, Jian Luan et al.KDD 2020 · 72 citations
Related papers
- SongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure BiasZihao Wang, Kejun Zhang, Yuxing Wang, Chen Zhang et al.ACM MM 2022 · 12 citations
- Structure-Enhanced Pop Music Generation via Harmony-Aware LearningXueyao Zhang, Jinchao Zhang, Yao Qiu, Li Wang et al.ACM MM 2022 · 24 citations
- PiRhDy: Learning Pitch-, Rhythm-, and Dynamics-aware Embeddings for Symbolic MusicHongru Liang, Wenqiang Lei, Paul Yaozhu Chan, Zhenglu Yang et al.ACM MM 2020 · 23 citations
- YuE: Scaling Open Foundation Models for Long-Form Music GenerationRuibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang et al.ICLR 2026 · 112 citations
- MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music CompositionPhilippe Pasquier, Jeff Ens, Nathan Fradet, Paul Triana et al.AAAI 2025 · 14 citations
