Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation
Botao Yu, Peiling Lu, Rui Wang, Wei Hu, Xu Tan, Wei Ye, Shikun Zhang, Tao Qin, Tie-Yan Liu
Abstract
Symbolic music generation aims to generate music scores automatically. A recent trend is to use Transformer or its variants in music generation, which is, however, suboptimal, because the full attention cannot efficiently model the typically long music sequences (e.g., over 10,000 tokens), and the existing models have shortcomings in generating musical repetition structures. In this paper, we propose Museformer, a Transformer with a novel fine-and coarse-grained attention for music generation. Specifically, with the fine-grained attention, a token of a specific bar directly attends to all the tokens of the bars that are most relevant to music structures (e.g., the previous 1st, 2nd, 4th and 8th bars, selected via similarity statistics); with the coarse-grained attention, a token only attends to the summarization of the other bars rather than each token of them so as to reduce the computational cost. The advantages are two-fold. First, it can capture both music structure-related correlations via the fine-grained attention, and other contextual information via the coarse-grained attention. Second, it is efficient and can model over 3× longer music sequences compared to its full-attention counterpart. Both objective and subjective experimental results demonstrate its ability to generate long music sequences with high quality and better structures. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Text2midi: Generating Symbolic Music from CaptionsKeshav Bhandari, Abhinaba Roy, Kyra Wang, Geeta Puri et al.AAAI 2025 · 21 citations
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale ProblemsSaeed Amizadeh, Sara Abdali, Yinheng Li, Kazuhito KoishidaNeurIPS 2025 · 2 citations
- Compose with Me: Collaborative Music Inpainter for Symbolic Music InfillingZhejing Hu, Yan Liu, Gong Chen, Bruce X. B. YuAAAI 2025 · 2 citations
- SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music GenerationHongrui Wang, Fan Zhang, Zhiyuan Yu, Ziya Zhou et al.ICLR 2026
- Diff-BGM: A Diffusion Model for Video Background Music GenerationSizhe Li, Yiming Qin, Minghang Zheng, Xin Jin et al.CVPR 2024
Builds on19
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- Random Feature AttentionHao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz et al.ICLR 2021 · 425 citations
Related papers
- N-gram Unsupervised Compoundation and Feature Injection for Better Symbolic Music UnderstandingJinhao Tian, Zuchao Li, Jiajia Li, Ping WangAAAI 2024
- FG-Midiformer: A Symbolic Music Understanding Model towards Fine-Grained Learning of Multi-AttributesHaonan Cheng, Junwei Zhang, Hengyan Huang, Long YeACM MM 2025 · 1 citation
- Encoding Musical Style with Transformer AutoencodersKristy Choi, Curtis Hawthorne, Ian Simon, Monica Dinculescu et al.ICML 2020 · 102 citations
- A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music ModelingZixun Guo, Jaeyong Kang, Dorien HerremansAAAI 2023 · 27 citations
- BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal StepsLekai Qian, Haoyu Gu, Jingwei Zhao, Ziyu WangICML 2026 · 3 citations
