MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch Estimation
Haojie Wei, Jun Yuan, Rui Zhang, Quanyu Dai, Yueguo Chen
Abstract
Music source separation and pitch estimation are two vital tasks in music information retrieval. Typically, the input of pitch estimation is obtained from the output of music source separation. Therefore, existing methods have tried to perform these two tasks simultaneously, so as to leverage the mutually beneficial relationship between both tasks. However, these methods still face two critical challenges that limit the improvement of both tasks: the lack of labeled data and joint learning optimization. To address these challenges, we propose a Model-Agnostic Joint Learning (MAJL) framework for both tasks. MAJL is a generic framework and can use variant models for each task. It includes a two-stage training method and a dynamic weighting method named Dynamic Weights on Hard Samples (DWHS), which addresses the lack of labeled data and joint learning optimization, respectively. Experimental results on public music datasets show that MAJL outperforms state-of-theart methods on both tasks, with significant improvements of 0.92 in Signal-to-Distortion Ratio (SDR) for music source separation and 2.71% in Raw Pitch Accuracy (RPA) for pitch estimation. Furthermore, comprehensive studies not only validate the effectiveness of each component of MAJL, but also indicate the great generality of MAJL in adapting to different model architectures.
• Applied computing → Sound and music computing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58a7687c-e305-4a05-b42d-c6eaa4ac5517Builds on5
- Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataKe Chen, Xingjian Du, Bilei Zhu, Zejun Ma et al.AAAI 2022 · 58 citations
- PMG : Personalized Multimodal Generation with Large Language ModelsXiaoteng Shen, Rui Zhang, Xiaoyan Zhao, Jieming Zhu et al.WWW 2024 · 40 citations
- Skipping the Frame-Level: Event-Based Piano Transcription With Neural Semi-CRFsYujia Yan, Frank Cwitkowitz, Zhiyao DuanNeurIPS 2021 · 32 citations
- Towards Multi-Intent Spoken Language Understanding via Hierarchical Attention and Optimal TransportXuxin Cheng, Zhihong Zhu, Hongxiang Li, Yaowei Li et al.AAAI 2024 · 19 citations
- MCSSME: Multi-Task Contrastive Learning for Semi-supervised Singing Melody Extraction from Polyphonic MusicShuai YuAAAI 2024 · 9 citations
Related papers
- Cyclic Co-Learning of Sounding Object Visual Grounding and Sound SeparationYapeng Tian, Di Hu, Chenliang XuCVPR 2021
- A Unified Audio-Visual Learning Framework for Localization, Separation, and RecognitionShentong Mo, Pedro MorgadoICML 2023 · 27 citations
- Modeling the Compatibility of Stem Tracks to Generate Music MashupsJiawen Huang, Ju-Chiang Wang, Jordan B. L. Smith, Xuchen Song et al.AAAI 2021 · 18 citations
- Weakly-supervised Audio Separation via Bi-modal Semantic SimilarityTanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida, Diana MarculescuICLR 2024 · 4 citations
- HKDSME: Heterogeneous Knowledge Distillation for Semi-supervised Singing Melody Extraction Using Harmonic SupervisionShuai Yu, Xiaoliang He, Ke Chen, Yi YuACM MM 2024 · 6 citations
