Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
Shitong Xu, Yiyuan Yang, Niki Trigoni, Andrew Markham
Abstract
Target speaker extraction focuses on isolating a specific speaker's voice from an audio mixture containing multiple speakers. To provide information about the target speaker's identity, prior works have utilized clean audio samples as conditioning inputs. However, such clean audio examples are not always readily available. For instance, obtaining a clean recording of a stranger's voice at a cocktail party without leaving the noisy environment is generally infeasible. Limited prior research has explored extracting the target speaker's characteristics from noisy enrollments, which may contain overlapping speech from interfering speakers. In this work, we explore a novel enrollment strategy that encodes target speaker information from the noisy enrollment by comparing segments where the target speaker is talking (Positive Enrollments) with segments where the target speaker is silent (Negative Enrollments). Experiments show the effectiveness of our model architecture, which achieves over 2.1 dB higher SI-SNRi compared to prior works in extracting the monaural speech from the mixture of two speakers. Additionally, the proposed two-stage training strategy accelerates convergence, reducing the number of optimization steps required to reach 3 dB SNR by 60%. Overall, our method achieves state-of-the-art performance in the monaural target speaker extraction conditioned on noisy enrollments. Our implementation is available at https://github.com/xu-shitong/TSE-through-Positive-Negative-Enroll .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3aab734f-91c6-4863-bba3-7ed4c2b1df4cBuilds on3
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- Look Once to Hear: Target Speech Hearing with Noisy ExamplesBandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka et al.CHI 2024 · 25 citations
- Self-Supervised Disentangled Representation Learning for Robust Target Speech ExtractionZhaoxi Mu, Xinyu Yang, Sining Sun, Qing YangAAAI 2024 · 13 citations
Related papers
- Trainable EEG Interpolation and Structure-Sharing Dual-Path Encoders for Brain-Assisted Target Speaker ExtractionZhao Lv, Haoran Zhou, Ying Chen, Youdian Gao et al.AAAI 2026
- Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker RepresentationsYejin Jeon, Yunsu Kim, Gary Geunbae LeeAAAI 2024 · 7 citations
- Tune-In: Training Under Negative Environments with Interference for Attention Networks Simulating Cocktail Party EffectJun Wang, Max W. Y. Lam, Dan Su, Dong YuAAAI 2021 · 7 citations
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao et al.ICLR 2021 · 64 citations
- Weakly-supervised Audio Separation via Bi-modal Semantic SimilarityTanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida, Diana MarculescuICLR 2024 · 4 citations
