N-gram Unsupervised Compoundation and Feature Injection for Better Symbolic Music Understanding
Jinhao Tian, Zuchao Li, Jiajia Li, Ping Wang
Abstract
The first step to apply deep learning techniques for symbolic music understanding is to transform musical pieces (mainly in MIDI format) into sequences of predefined tokens like note pitch, note velocity, and chords. Subsequently, the sequences are fed into a neural sequence model to accomplish specific tasks. Music sequences exhibit strong correlations between adjacent elements, making them prime candidates for N-gram techniques from Natural Language Processing (NLP). Consider classical piano music: specific melodies might recur throughout a piece, with subtle variations each time. In this paper, we propose a novel method, NG-Midiformer, for understanding symbolic music sequences that leverages the Ngram approach. Our method involves first processing music pieces into word-like sequences with our proposed unsupervised compoundation, followed by using our N-gram Transformer encoder, which can effectively incorporate N-gram information to enhance the primary encoder part for better understanding of music sequences. The pre-training process on large-scale music datasets enables the model to thoroughly learn the N-gram information contained within music sequences, and subsequently apply this information for making inferences during the fine-tuning stage. Experiment on various datasets demonstrate the effectiveness of our method and achieved state-of-the-art performance on a series of music understanding downstream tasks. The code and model weights will be released at https://github.com/CinqueOrigin/ NG-Midiformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years AwayJiliang Hu, Jiajia Li, Ziyi Pan, Chong Chen et al.AAAI 2025 · 3 citations
- Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion RecognitionHaiying Xia, Zhongyi Huang, Yumei Tan, Shuxiang SongAAAI 2026
Builds on3
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsYu-Siang Huang, Yi-Hsuan YangACM MM 2020 · 265 citations
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed HypergraphsWen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, Yi-Hsuan YangAAAI 2021 · 242 citations
Related papers
- FG-Midiformer: A Symbolic Music Understanding Model towards Fine-Grained Learning of Multi-AttributesHaonan Cheng, Junwei Zhang, Hengyan Huang, Long YeACM MM 2025 · 1 citation
- Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationBotao Yu, Peiling Lu, Rui Wang, Wei Hu et al.NeurIPS 2022 · 104 citations
- Byte Pair Encoding for Symbolic MusicNathan Fradet, Nicolas Gutowski, Fabien Chhel, Jean-Pierre BriotEMNLP 2023 · 8 citations
- A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music ModelingZixun Guo, Jaeyong Kang, Dorien HerremansAAAI 2023 · 27 citations
- Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-TrainingHong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia et al.ICML 2026 · 3 citations
