Improving Target Sound Extraction via Disentangled Codec Representations with Privileged Knowledge Distillation
Dail Kim, Joon-Hyuk Chang
摘要
Target sound extraction aims to isolate target sound sources from an input mixture using a target clue to identify the sounds of interest. To address the challenge posed by the wide variety of sounds, recent work has introduced privileged knowledge distillation (PKD), which utilizes privileged information (PI) about the target sound, available only during training. While PKD has shown promise, existing approaches often suffer from overfitting of the teacher model for the overly rich PI and ineffective knowledge transfer to the student model. In this paper, we propose Disentangled Codec Knowledge Distillation (DCKD) to mitigate these issues by regulating the amount and the flow of target sound information within the teacher model. We begin by extracting a compressed representation of the target sound using a neural audio codec to regulate the amount of PI. Disentangled representation learning is then applied to remove class information and extract fine-grained temporal information as PI. Subsequently, an n-hot vector as the class information and the class-independent PI are used to condition the early and later layers of the teacher model, respectively, forming a regulated coarse-to-fine target information flow. The resulting representation is transferred to the student model through feature-level knowledge distillation. Experimental results show that DCKD consistently improves existing methods across model architectures under the multi-target selection condition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 等ICML 2020 · 被引用 512 次
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersKaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni 等ICML 2022 · 被引用 157 次
- Contrastively Disentangled Sequential Variational AutoencoderJunwen Bai, Weiran Wang, Carla P. GomesNeurIPS 2021 · 被引用 60 次
- Toward Understanding Privileged Features Distillation in Learning-to-RankShuo Yang, Sujay Sanghavi, Holakou Rahmanian, Jan Bakus 等NeurIPS 2022 · 被引用 31 次
相关 Paper
- Prototypical Contrastive Predictive CodingKyungmin LeeICLR 2022 · 被引用 9 次
- Wasserstein Contrastive Representation DistillationLiqun Chen, Dong Wang, Zhe Gan, Jingjing Liu 等CVPR 2021
- Entropy-Monitored Kernelized Token Distillation for Audio-Visual CompressionHyoungseob Park, Lipeng Ke, Pritish Mohapatra, Huajun Ying 等ICLR 2026
- Comprehensive Knowledge Distillation with Causal InterventionXiang Deng, Zhongfei ZhangNeurIPS 2021 · 被引用 44 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
