MIDI-Zero: A MIDI-driven Self-Supervised Learning Approach for Music Retrieval
Yuhang Su, Wei Hu, Hongfeng Gao, Fan Zhang
摘要
Content-based Music Retrieval (CBMR) is a fundamental task in music information retrieval, encompassing sub-tasks including Audio Identification, Audio Matching, and Version Identification. Traditional methods typically analyze audio signals or spectrograms to extract features related to rhythm, melody, harmony, and timbre. However, with the rapid development of Music Transcription and digital music technologies, MIDI representation has emerged as a powerful alternative fo r music analysis. In this paper, we propose MIDI-Zero, a novel self-supervisedlearning framework for CBMR that operates entirely on MIDI representations. Unlike existing approaches, MIDI-Zero requires no external training data; all training data is automatically generated based on predefined task rules, eliminating the need for labeled datasets or external music collections. MIDI-Zero is designed to handle both symbolic music data and audio-based tasks by leveraging Music Transcription models. Its strong robustness ensures effectiveness even with low-quality transcriptions. Extensive experiments demonstrate that MIDI-Zero achieves competitive performance across various CBMR sub-tasks, particularly excelling in Audio Matching. Our approach simplifies the feature extraction process, bridges the gap between audio and symbolic music representations, and offers a versatile and scalable solution for music retrieval.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- MusicBERT: A Self-supervised Learning of Music RepresentationHongyuan Zhu, Ye Niu, Di Fu, Hao WangACM MM 2021 · 被引用 17 次
- MCSSME: Multi-Task Contrastive Learning for Semi-supervised Singing Melody Extraction from Polyphonic MusicShuai YuAAAI 2024 · 被引用 9 次
- Contrastive Learning with Positive-Negative Frame Mask for Music RepresentationDong Yao, Zhou Zhao, Shengyu Zhang, Jieming Zhu 等WWW 2022 · 被引用 26 次
- Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured TokenizationLongshen Ou, Jingwei Zhao, Ziyu Wang, Gus Xia 等NeurIPS 2025 · 被引用 5 次
- Unaligned Supervision for Automatic Music Transcription in The WildBen Maman, Amit H. BermanoICML 2022 · 被引用 46 次
