FG-Midiformer: A Symbolic Music Understanding Model towards Fine-Grained Learning of Multi-Attributes
Haonan Cheng, Junwei Zhang, Hengyan Huang, Long Ye
摘要
Symbolic music understanding is a fundamental task in multimedia interpretation, which aims to decode musical attributes from symbolic music representations. Existing methods usually handle musical sequences as linguistic data, ignoring intrinsic musical properties. For instance, recurring melodies often appear in musical pieces with subtle variations, which requires methods that are aware of both local musical details and overall repetitive patterns, yet current methods can not fit the request. To address this issue, we introduce FG-Midiformer, a modified transformer framework that incorporates a multi-scale-aware feature learning (MSAFL) module and a local feature enhanced classification (LFEC) module for fine-grained understanding of multi-attributes. Specifically, the MSAFL module is designed to capture multi-scale musical relationships by embedding efficient multi-scale attention for long-term dependency modeling. In order to improve the classification accuracy of musical attributes, we devise the LFEC module, in which an attention mechanism with full 3-D weights is first introduced to efficiently highlight and leverage important local musical features. The LFEC module strengthens local feature representation and improves the sensitivity of FG-Midiformer to subtle differences between musical attributes. Extensive experiments show that FG-Midiformer achieves state-of-the-art performance in multi-attribute understanding tasks such as melody identification, velocity prediction, composer categorization, and emotion classification. The code will be released at https://github.com/Viki66666/FG-Midiformer.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- N-gram Unsupervised Compoundation and Feature Injection for Better Symbolic Music UnderstandingJinhao Tian, Zuchao Li, Jiajia Li, Ping WangAAAI 2024
- Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationBotao Yu, Peiling Lu, Rui Wang, Wei Hu 等NeurIPS 2022 · 被引用 104 次
- A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music ModelingZixun Guo, Jaeyong Kang, Dorien HerremansAAAI 2023 · 被引用 27 次
- Fine-Grained Position Helps Memorizing More, a Novel Music Compound Transformer Model with Feature Interaction FusionZuchao Li, Ruhan Gong, Yineng Chen, Kehua SuAAAI 2023 · 被引用 11 次
- Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion RecognitionHaiying Xia, Zhongyi Huang, Yumei Tan, Shuxiang SongAAAI 2026
