Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction
Zhaoxi Mu, Xinyu Yang, Sining Sun, Qing Yang
摘要
Speech signals are inherently complex as they encompass both global acoustic characteristics and local semantic information. However, in the task of target speech extraction, certain elements of global and local semantic information in the reference speech, which are irrelevant to speaker identity, can lead to speaker confusion within the speech extraction network. To overcome this challenge, we propose a self-supervised disentangled representation learning method. Our approach tackles this issue through a two-phase process, utilizing a reference speech encoding network and a global information disentanglement network to gradually disentangle the speaker identity information from other irrelevant factors. We exclusively employ the disentangled speaker identity information to guide the speech extraction network. Moreover, we introduce the adaptive modulation Transformer to ensure that the acoustic representation of the mixed signal remains undisturbed by the speaker embeddings. This component incorporates speaker embeddings as conditional information, facilitating natural and efficient guidance for the speech extraction network. Experimental results substantiate the effectiveness of our meticulously crafted approach, showcasing a substantial reduction in the likelihood of speaker confusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Target Speaker Extraction through Comparing Noisy Positive and Negative Audio EnrollmentsShitong Xu, Yiyuan Yang, Niki Trigoni, Andrew MarkhamNeurIPS 2025 · 被引用 4 次
- Improving Target Sound Extraction via Disentangled Codec Representations with Privileged Knowledge DistillationDail Kim, Joon-Hyuk ChangNeurIPS 2025 · 被引用 1 次
- MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor DisentanglementXinyue Yu, Youqing Fang, Pingyu Wu, Guoyang Ye 等AAAI 2026
它引用的顶会 Paper4
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 等ICML 2020 · 被引用 512 次
- Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech GenerationDongchan Min, Dong Bok Lee, Eunho Yang, Sung Ju HwangICML 2021 · 被引用 218 次
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersKaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni 等ICML 2022 · 被引用 157 次
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
相关 Paper
- Disentangling Voice and Content with Self-Supervision for Speaker RecognitionTianchi Liu, Kong Aik Lee, Qiongqiong Wang, Haizhou LiNeurIPS 2023 · 被引用 53 次
- VoiceMixer: Adversarial Voice Style MixupSang-Hoon Lee, Ji-Hoon Kim, Hyunseung Chung, Seong-Whan LeeNeurIPS 2021 · 被引用 46 次
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao 等ICLR 2021 · 被引用 64 次
- Cross-modal Self-Supervised Learning for Lip Reading: When Contrastive Learning meets Adversarial TrainingChangchong Sheng, Matti Pietikäinen, Qi Tian, Li LiuACM MM 2021 · 被引用 11 次
- Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker RepresentationsYejin Jeon, Yunsu Kim, Gary Geunbae LeeAAAI 2024 · 被引用 7 次
