MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch Estimation
Haojie Wei, Jun Yuan, Rui Zhang, Quanyu Dai, Yueguo Chen
摘要
Music source separation and pitch estimation are two vital tasks in music information retrieval. Typically, the input of pitch estimation is obtained from the output of music source separation. Therefore, existing methods have tried to perform these two tasks simultaneously, so as to leverage the mutually beneficial relationship between both tasks. However, these methods still face two critical challenges that limit the improvement of both tasks: the lack of labeled data and joint learning optimization. To address these challenges, we propose a Model-Agnostic Joint Learning (MAJL) framework for both tasks. MAJL is a generic framework and can use variant models for each task. It includes a two-stage training method and a dynamic weighting method named Dynamic Weights on Hard Samples (DWHS), which addresses the lack of labeled data and joint learning optimization, respectively. Experimental results on public music datasets show that MAJL outperforms state-of-theart methods on both tasks, with significant improvements of 0.92 in Signal-to-Distortion Ratio (SDR) for music source separation and 2.71% in Raw Pitch Accuracy (RPA) for pitch estimation. Furthermore, comprehensive studies not only validate the effectiveness of each component of MAJL, but also indicate the great generality of MAJL in adapting to different model architectures.
• Applied computing → Sound and music computing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataKe Chen, Xingjian Du, Bilei Zhu, Zejun Ma 等AAAI 2022 · 被引用 58 次
- PMG : Personalized Multimodal Generation with Large Language ModelsXiaoteng Shen, Rui Zhang, Xiaoyan Zhao, Jieming Zhu 等WWW 2024 · 被引用 40 次
- Skipping the Frame-Level: Event-Based Piano Transcription With Neural Semi-CRFsYujia Yan, Frank Cwitkowitz, Zhiyao DuanNeurIPS 2021 · 被引用 32 次
- Towards Multi-Intent Spoken Language Understanding via Hierarchical Attention and Optimal TransportXuxin Cheng, Zhihong Zhu, Hongxiang Li, Yaowei Li 等AAAI 2024 · 被引用 19 次
- MCSSME: Multi-Task Contrastive Learning for Semi-supervised Singing Melody Extraction from Polyphonic MusicShuai YuAAAI 2024 · 被引用 9 次
相关 Paper
- Cyclic Co-Learning of Sounding Object Visual Grounding and Sound SeparationYapeng Tian, Di Hu, Chenliang XuCVPR 2021
- A Unified Audio-Visual Learning Framework for Localization, Separation, and RecognitionShentong Mo, Pedro MorgadoICML 2023 · 被引用 27 次
- Modeling the Compatibility of Stem Tracks to Generate Music MashupsJiawen Huang, Ju-Chiang Wang, Jordan B. L. Smith, Xuchen Song 等AAAI 2021 · 被引用 18 次
- Weakly-supervised Audio Separation via Bi-modal Semantic SimilarityTanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida, Diana MarculescuICLR 2024 · 被引用 4 次
- HKDSME: Heterogeneous Knowledge Distillation for Semi-supervised Singing Melody Extraction Using Harmonic SupervisionShuai Yu, Xiaoliang He, Ke Chen, Yi YuACM MM 2024 · 被引用 6 次
