MusicBERT: A Self-supervised Learning of Music Representation
Hongyuan Zhu, Ye Niu, Di Fu, Hao Wang
Abstract
Music recommendation has been one of the most used information retrieval services on internet. Finding suitable music for users' demands from tens of millions of music relies on the understanding of music content. Traditional studies usually focus on music representation based on massive user behavioral data and music meta-data, which ignore the audio characteristic of music. However, it is found that the melodic characteristics of music themselves can be further used to understand music. Moreover, how to utilize large-scale audio data to learn music representation is not well explored. To this end, we propose a self-supervised learning model for music representation. We firstly utilize a beat-level music pre-training model to learn the structure of music. Then, we use a multi-task learning framework to model music self-representation and co-relations between music, concurrently. Besides, we propose several downstream tasks to evaluate music representation, including music genre classification, music highlight, and music similarity retrieval. Extensive experiments on multiple music datasets demonstrate our model's superiority over baselines on learning music representation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7e2c1e2d-1fb6-4ef5-b755-79255de48728Cited by top-tier papers6
- MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingYizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma et al.ICLR 2024 · 277 citations
- A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music ModelingZixun Guo, Jaeyong Kang, Dorien HerremansAAAI 2023 · 27 citations
- Mimicking the Annotation Process for Recognizing the Micro ExpressionsBo-Kai Ruan, Ling Lo, Hong-Han Shuai, Wen-Huang ChengACM MM 2022 · 16 citations
- MCSSME: Multi-Task Contrastive Learning for Semi-supervised Singing Melody Extraction from Polyphonic MusicShuai YuAAAI 2024 · 9 citations
- DisCover: Disentangled Music Representation Learning for Cover Song IdentificationJiahao Xun, Shengyu Zhang, Yanting Yang, Jieming Zhu et al.SIGIR 2023 · 7 citations
Related papers
- MIDI-Zero: A MIDI-driven Self-Supervised Learning Approach for Music RetrievalYuhang Su, Wei Hu, Hongfeng Gao, Fan ZhangSIGIR 2025
- Exploiting Behavioral Consistence for Universal User RepresentationJie Gu, Feng Wang, Qinghui Sun, Zhiquan Ye et al.AAAI 2021 · 32 citations
- AudioMosaic: Contrastive Masked Audio Representation LearningHanxun Huang, Qizhou Wang, Xingjun Ma, Cihang Xie et al.ICML 2026 · 2 citations
- Contrastive Learning with Positive-Negative Frame Mask for Music RepresentationDong Yao, Zhou Zhao, Shengyu Zhang, Jieming Zhu et al.WWW 2022 · 26 citations
- It's Time for Artistic Correspondence in Music and VideoDídac Surís, Carl Vondrick, Bryan C. Russell, Justin SalamonCVPR 2022 · 33 citations
