MCSSME: Multi-Task Contrastive Learning for Semi-supervised Singing Melody Extraction from Polyphonic Music
Shuai Yu
Abstract
Singing melody extraction is an important task in the field of music information retrieval (MIR). The development of data-driven models for this task have achieved great successes. However, the existing models have two major limitations: firstly, most of the existing singing melody extraction models have formulated this task as a pixel-level prediction task. The lack of labeling data has limited the model for further improvements. Secondly, the generalization of the existing models are prone to be disturbed by the music genres. To address the issues mentioned above, in this paper, we propose a multi-Task contrastive learning framework for semi-supervised singing melody extraction, termed as MCSSME. Specifically, to deal with data scarcity limitation, we propose a self-consistency regularization (SCR) method to train the model on the unlabeled data. Transformations are applied to the raw signal of polyphonic music, which makes the network to improve its representation capability via recognizing the transformations. We further propose a novel multi-task learning (MTL) approach to jointly learn singing melody extraction and classification of transformed data. To deal with generalization limitation, we also propose a contrastive embedding learning, which strengthens the intra-class compactness and inter-class separability. To improve the generalization on different music genres, we also propose a domain classification method to learn task-dependent features by mapping data from different music genres to shared subspace. MCSSME evaluates on a set of well-known public melody extraction datasets with promising performances. The experimental results demonstrate the effectiveness of the MCSSME framework for singing melody extraction from polyphonic music using very limited labeled data scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbb7fd49-e655-4f18-bd3c-53e7508c2814Cited by top-tier papers3
- MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch EstimationHaojie Wei, Jun Yuan, Rui Zhang, Quanyu Dai et al.ACM MM 2024 · 3 citations
- Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion ModelShuai Yu, Xiaoliang He, Kangjie Dong, Yi YuAAAI 2026
- Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMsZhejing Hu, Yan Liu, Zhi Zhang, Aiwei Zhang et al.AAAI 2026
Builds on7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- MusicBERT: A Self-supervised Learning of Music RepresentationHongyuan Zhu, Ye Niu, Di Fu, Hao WangACM MM 2021 · 17 citations
Related papers
- HKDSME: Heterogeneous Knowledge Distillation for Semi-supervised Singing Melody Extraction Using Harmonic SupervisionShuai Yu, Xiaoliang He, Ke Chen, Yi YuACM MM 2024 · 6 citations
- DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic MusicShuai Yu, Xiaoliang He, Kangjie Dong, Yi YuACM MM 2025 · 1 citation
- Contrastive Learning with Positive-Negative Frame Mask for Music RepresentationDong Yao, Zhou Zhao, Shengyu Zhang, Jieming Zhu et al.WWW 2022 · 26 citations
- Partial Multi-View Clustering via Self-Supervised NetworkWei Feng, Guoshuai Sheng, Qianqian Wang, Quanxue Gao et al.AAAI 2024 · 15 citations
- MIDI-Zero: A MIDI-driven Self-Supervised Learning Approach for Music RetrievalYuhang Su, Wei Hu, Hongfeng Gao, Fan ZhangSIGIR 2025
