MuPT: A Generative Symbolic Music Pretrained Transformer
Xingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou, Ka Man Lo, Jiaheng Liu, Ruibin Yuan, Lejun Min, Xueling Liu, Tianyu Zhang, Xeron Du, Shuyue Guo
Abstract
In this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition. To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a Synchronized Multi-Track ABC Notation (SMT-ABC Notation), which aims to preserve coherence across multiple musical tracks. Our contributions include a series of models capable of handling up to 8192 tokens, covering 90% of the symbolic music data in our training set. Furthermore, we explore the implications of the Symbolic Music Scaling Law (SMS Law) on model performance. The results indicate a promising direction for future research in music generation, offering extensive resources for community-led research through our open-source contributions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language ModelsHaoran Que, Jiaheng Liu, Ge Zhang, Chenchen Zhang et al.NeurIPS 2024 · 47 citations
- Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic MusicHongju Su, Ke Li, Lan Yang, Honggang Zhang et al.ACL 2026
- Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMsZhejing Hu, Yan Liu, Zhi Zhang, Aiwei Zhang et al.AAAI 2026
Builds on12
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu et al.ICML 2022 · 1,123 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
Related papers
- Text2midi: Generating Symbolic Music from CaptionsKeshav Bhandari, Abhinaba Roy, Kyra Wang, Geeta Puri et al.AAAI 2025 · 21 citations
- SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-TrainingJiaxing Yu, Xinda Wu, Yunfei Xu, Tieyao Zhang et al.AAAI 2025 · 2 citations
- SongComposer: A Large Language Model for Lyric and Melody Generation in Song CompositionShuangrui Ding, Zihan Liu, Xiaoyi Dong, Pan Zhang et al.ACL 2025
- Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical ScoresCongren Dai, Yue Yang, Krinos Li, Huichi Zhou et al.ACL 2026 · 4 citations
- N-gram Unsupervised Compoundation and Feature Injection for Better Symbolic Music UnderstandingJinhao Tian, Zuchao Li, Jiajia Li, Ping WangAAAI 2024
