Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion Model
Shuai Yu, Xiaoliang He, Kangjie Dong, Yi Yu
Abstract
Semi-supervised singing melody extraction (SSME) is one of the key tasks in the field of music information retrieval (MIR). Recently, several SSME methods have been proposed and achieved remarkable successes. However, existing methods are still facing two critical issues: firstly, there is a lack of an effective data augmentation method for SSME, which results in insufficient utilization of unlabeled data. Secondly, existing SSME methods discards too much unlabeled data in the stage of consistency regularization, which hinders the further improvements of SSME task. In this paper, we present ELH-SME, a novel framework that better utilizes the unlabeled musical data for SSME task. Specifically, our proposed ELH-SME framework consists of three modules: (1) we first propose a diffusion-based multi-bands augmentation (DMA) method to increase the amounts of training data. The proposed DMA methods employs a diffusion model to generate perturbation at the specific frequency bands in an end-to-end manner, thereby avoiding sharply perturbations to the spectrogram. (2) To improve the utilization rate of unlabeled data, we suggest a global-class confidence (GCC) module. During the phase of consistency regularization, we consider both the global-wise and class-wise confidence values, improving the utilization rate of unlabeled data. (3) To further improve the utilization of unlabeled data, we also propose to enhance the representation capability of unlabeled data by extracting channel-level features from labeled data via channel cross attention (CCA). We evaluate our proposed framework on several well-known public available datasets, and the conducted experiments demonstrate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7d4124b-82e6-41d0-a691-a8fbc97dcb16Builds on17
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Fast Timing-Conditioned Latent Audio DiffusionZach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley et al.ICML 2024 · 220 citations
- Consistency Regularization for Certified Robustness of Smoothed ClassifiersJongheon Jeong, Jinwoo ShinNeurIPS 2020 · 103 citations
- OpenMatch: Open-Set Semi-supervised Learning with Open-set Consistency RegularizationKuniaki Saito, Donghyun Kim, Kate SaenkoNeurIPS 2021 · 80 citations
Related papers
- HKDSME: Heterogeneous Knowledge Distillation for Semi-supervised Singing Melody Extraction Using Harmonic SupervisionShuai Yu, Xiaoliang He, Ke Chen, Yi YuACM MM 2024 · 6 citations
- DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic MusicShuai Yu, Xiaoliang He, Kangjie Dong, Yi YuACM MM 2025 · 1 citation
- MCSSME: Multi-Task Contrastive Learning for Semi-supervised Singing Melody Extraction from Polyphonic MusicShuai YuAAAI 2024 · 9 citations
- Complex-Cycle-Consistent Diffusion Model for Monaural Speech EnhancementYi Li, Yang Sun, Plamen P. AngelovAAAI 2025 · 2 citations
- Multi-view Self-Supervised Contrastive Learning for Multivariate Time SeriesYuhan Wu, Xiyu Meng, Yang He, Junru Zhang et al.ACM MM 2024 · 5 citations
