PopMAG: Pop Music Accompaniment Generation
Yi Ren, Jinzheng He, Xu Tan, Tao Qin, Zhou Zhao, Tie-Yan Liu
摘要
In pop music, accompaniments are usually played by multiple instruments (tracks) such as drum, bass, string and guitar, and can make a song more expressive and contagious by arranging together with its melody. Previous works usually generate multiple tracks separately and the music notes from different tracks not explicitly depend on each other, which hurts the harmony modeling. To improve harmony, in this paper 1 , we propose a novel MUlti-track MIDI representation (MuMIDI), which enables simultaneous multi-track generation in a single sequence and explicitly models the dependency of the notes from different tracks. While this greatly improves harmony, unfortunately, it enlarges the sequence length and brings the new challenge of long-term music modeling. We further introduce two new techniques to address this challenge: 1) We model multiple note attributes (e.g., pitch, duration, velocity) of a musical note in one step instead of multiple steps, which can shorten the length of a MuMIDI sequence. 2) We introduce extra long-context as memory to capture long-term dependency in music. We call our system for pop music accompaniment generation as PopMAG. We evaluate PopMAG on multiple datasets (LMD, FreeMidi and CPMD, a private dataset of Chinese pop songs) with both subjective and objective metrics. The results demonstrate the effectiveness of PopMAG for multi-track harmony modeling and long-term context modeling. Specifically, PopMAG wins 42%/38%/40% votes when comparing with ground truth musical pieces on LMD, FreeMidi and CPMD datasets respectively and largely outperforms other state-ofthe-art music accompaniment generation models and multi-track MIDI representations in terms of subjective and objective metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationBotao Yu, Peiling Lu, Rui Wang, Wei Hu 等NeurIPS 2022 · 被引用 104 次
- Video Background Music Generation with Controllable Music TransformerShangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang 等ACM MM 2021 · 被引用 87 次
- SongMASS: Automatic Song Writing with Pre-training and Alignment ConstraintZhonghao Sheng, Kaitao Song, Xu Tan, Yi Ren 等AAAI 2021 · 被引用 84 次
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao 等ICCV 2023 · 被引用 51 次
- Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion ModelsZiyu Wang, Lejun Min, Gus XiaICLR 2024 · 被引用 32 次
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsYu-Siang Huang, Yi-Hsuan YangACM MM 2020 · 被引用 265 次
- Encoding Musical Style with Transformer AutoencodersKristy Choi, Curtis Hawthorne, Ian Simon, Monica Dinculescu 等ICML 2020 · 被引用 102 次
- DeepSinger: Singing Voice Synthesis with Data Mined From the WebYi Ren, Xu Tan, Tao Qin, Jian Luan 等KDD 2020 · 被引用 72 次
相关 Paper
- SongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure BiasZihao Wang, Kejun Zhang, Yuxing Wang, Chen Zhang 等ACM MM 2022 · 被引用 12 次
- Structure-Enhanced Pop Music Generation via Harmony-Aware LearningXueyao Zhang, Jinchao Zhang, Yao Qiu, Li Wang 等ACM MM 2022 · 被引用 24 次
- PiRhDy: Learning Pitch-, Rhythm-, and Dynamics-aware Embeddings for Symbolic MusicHongru Liang, Wenqiang Lei, Paul Yaozhu Chan, Zhenglu Yang 等ACM MM 2020 · 被引用 23 次
- YuE: Scaling Open Foundation Models for Long-Form Music GenerationRuibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang 等ICLR 2026 · 被引用 112 次
- MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music CompositionPhilippe Pasquier, Jeff Ens, Nathan Fradet, Paul Triana 等AAAI 2025 · 被引用 14 次
