Fine-Grained Position Helps Memorizing More, a Novel Music Compound Transformer Model with Feature Interaction Fusion
Zuchao Li, Ruhan Gong, Yineng Chen, Kehua Su
Abstract
Due to the particularity of the simultaneous occurrence of multiple events in music sequences, compound Transformer is proposed to deal with the challenge of long sequences. However, there are two deficiencies in the compound Transformer. First, since the order of events is more important for music than natural language, the information provided by the original absolute position embedding is not precise enough. Second, there is an important correlation between the tokens in the compound word, which is ignored by the current compound Transformer. Therefore, in this work, we propose an improved compound Transformer model for music understanding. Specifically, we propose an attribute embedding fusion module and a novel position encoding scheme with absolute-relative consideration. In the attribute embedding fusion module, different attributes are fused through feature permutation by using a multi-head self-attention mechanism in order to capture rich interactions between attributes. In the novel position encoding scheme, we propose RoAR position encoding, which realizes rotational absolute position encoding, relative position encoding, and absolute-relative position interactive encoding, providing clear and rich orders for musical events. Empirical study on four typical music understanding tasks shows that our attribute fusion approach and RoAR position encoding brings large performance gains. In addition, we further investigate the impact of masked language modeling and casual language modeling pre-training on music understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9dbc348-3a6f-457d-a993-bb83fa7e380fBuilds on9
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- InteractE: Improving Convolution-Based Knowledge Graph Embeddings by Increasing Feature InteractionsShikhar Vashishth, Soumya Sanyal, Vikram Nitin, Nilesh Agrawal et al.AAAI 2020 · 393 citations
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 358 citations
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsYu-Siang Huang, Yi-Hsuan YangACM MM 2020 · 265 citations
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed HypergraphsWen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, Yi-Hsuan YangAAAI 2021 · 242 citations
Related papers
- A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music ModelingZixun Guo, Jaeyong Kang, Dorien HerremansAAAI 2023 · 27 citations
- Decoupling The "What" and "Where" With Polar Coordinate Positional EmbeddingAnand Gopalakrishnan, Róbert Csordás, Jürgen Schmidhuber, Michael MozerICML 2026 · 7 citations
- FG-Midiformer: A Symbolic Music Understanding Model towards Fine-Grained Learning of Multi-AttributesHaonan Cheng, Junwei Zhang, Hengyan Huang, Long YeACM MM 2025 · 1 citation
- HiRoPE: Length Extrapolation for Code Models Using Hierarchical PositionKechi Zhang, Ge Li, Huangzhao Zhang, Zhi JinACL 2024
- PermuteFormer: Efficient Relative Position Encoding for Long SequencesPeng ChenEMNLP 2021 · 16 citations
