Improving Target Sound Extraction via Disentangled Codec Representations with Privileged Knowledge Distillation
Dail Kim, Joon-Hyuk Chang
Abstract
Target sound extraction aims to isolate target sound sources from an input mixture using a target clue to identify the sounds of interest. To address the challenge posed by the wide variety of sounds, recent work has introduced privileged knowledge distillation (PKD), which utilizes privileged information (PI) about the target sound, available only during training. While PKD has shown promise, existing approaches often suffer from overfitting of the teacher model for the overly rich PI and ineffective knowledge transfer to the student model. In this paper, we propose Disentangled Codec Knowledge Distillation (DCKD) to mitigate these issues by regulating the amount and the flow of target sound information within the teacher model. We begin by extracting a compressed representation of the target sound using a neural audio codec to regulate the amount of PI. Disentangled representation learning is then applied to remove class information and extract fine-grained temporal information as PI. Subsequently, an n-hot vector as the class information and the class-independent PI are used to condition the early and later layers of the teacher model, respectively, forming a regulated coarse-to-fine target information flow. The resulting representation is transferred to the student model through feature-level knowledge distillation. Experimental results show that DCKD consistently improves existing methods across model architectures under the multi-target selection condition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89fa54a6-32bb-41c3-9fd3-065780c8f2abBuilds on7
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar et al.NeurIPS 2023 · 910 citations
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersKaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni et al.ICML 2022 · 157 citations
- Contrastively Disentangled Sequential Variational AutoencoderJunwen Bai, Weiran Wang, Carla P. GomesNeurIPS 2021 · 60 citations
- Toward Understanding Privileged Features Distillation in Learning-to-RankShuo Yang, Sujay Sanghavi, Holakou Rahmanian, Jan Bakus et al.NeurIPS 2022 · 31 citations
Related papers
- Prototypical Contrastive Predictive CodingKyungmin LeeICLR 2022 · 9 citations
- Wasserstein Contrastive Representation DistillationLiqun Chen, Dong Wang, Zhe Gan, Jingjing Liu et al.CVPR 2021
- Entropy-Monitored Kernelized Token Distillation for Audio-Visual CompressionHyoungseob Park, Lipeng Ke, Pritish Mohapatra, Huajun Ying et al.ICLR 2026
- Comprehensive Knowledge Distillation with Causal InterventionXiang Deng, Zhongfei ZhangNeurIPS 2021 · 44 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
