Contrastive Learning with Positive-Negative Frame Mask for Music Representation
Dong Yao, Zhou Zhao, Shengyu Zhang, Jieming Zhu, Yudong Zhu, Rui Zhang, Xiuqiang He
Abstract
Self-supervised learning, especially contrastive learning, has made an outstanding contribution to the development of many deep learning research fields. Recently, researchers in the acoustic signal processing field noticed its success and leveraged contrastive learning for better music representation. Typically, existing approaches maximize the similarity between two distorted audio segments sampled from the same music. In other words, they ensure a semantic agreement at the music level. However, those coarse-grained methods neglect some inessential or noisy elements at the frame level, which may be detrimental to the model to learn the effective representation of music. Towards this end, this paper proposes a novel Positive-nEgative frame mask for Music Representation based on the contrastive learning framework, abbreviated as PEMR. Concretely, PEMR incorporates a Positive-Negative Mask Generation module, which leverages transformer blocks to generate frame masks on Log-Mel spectrogram. We can generate self-augmented negative and positive samples by masking important components or inessential components, respectively. We devise a novel contrastive learning objective to accommodate both self-augmented positives/negatives sampled from the same music. We conduct experiments on four public datasets. The experimental results of two music-related downstream tasks, music classification and cover song identification, demonstrate the generalization ability and transferability of music representation learned by PEMR 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 279781f6-0229-415c-8b4c-d14454980190Cited by top-tier papers4
- Contrastive Balancing Representation Learning for Heterogeneous Dose-Response Curves EstimationMinqin Zhu, Anpeng Wu, Haoxuan Li, Ruoxuan Xiong et al.AAAI 2024 · 12 citations
- Hierarchical Topology Isomorphism Expertise Embedded Graph Contrastive LearningJiangmeng Li, Yifan Jin, Hang Gao, Wenwen Qiang et al.AAAI 2024 · 10 citations
- DisCover: Disentangled Music Representation Learning for Cover Song IdentificationJiahao Xun, Shengyu Zhang, Yanting Yang, Jieming Zhu et al.SIGIR 2023 · 7 citations
- Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMsZhejing Hu, Yan Liu, Zhi Zhang, Aiwei Zhang et al.AAAI 2026
Builds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
Related papers
- MCSSME: Multi-Task Contrastive Learning for Semi-supervised Singing Melody Extraction from Polyphonic MusicShuai YuAAAI 2024 · 9 citations
- AudioMosaic: Contrastive Masked Audio Representation LearningHanxun Huang, Qizhou Wang, Xingjun Ma, Cihang Xie et al.ICML 2026 · 2 citations
- Language Pre-training Guided Masking Representation Learning for Time Series ClassificationLiaoyuan Tang, Zheng Wang, Jie Wang, Guanxiong He et al.AAAI 2025 · 1 citation
- Contrastive Audio-Visual Masked AutoencoderYuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath et al.ICLR 2023 · 17 citations
- PointCMP: Contrastive Mask Prediction for Self-supervised Learning on Point Cloud VideosZhiqiang Shen, Xiaoxiao Sheng, Longguang Wang, Yulan Guo et al.CVPR 2023
