UNSSOR: Unsupervised Neural Speech Separation by Leveraging Over-determined Training Mixtures
Zhong-Qiu Wang, Shinji Watanabe
Abstract
In reverberant conditions with multiple concurrent speakers, each microphone acquires a mixture signal of multiple speakers at a different location. In overdetermined conditions where the microphones out-number speakers, we can narrow down the solutions to speaker images and realize unsupervised speech separation by leveraging each mixture signal as a constraint (i.e., the estimated speaker images at a microphone should add up to the mixture). Equipped with this insight, we propose UNSSOR, an algorithm for unsupervised neural speech separation by leveraging over-determined training mixtures. At each training step, we feed an input mixture to a deep neural network (DNN) to produce an intermediate estimate for each speaker, linearly filter the estimates, and optimize a loss so that, at each microphone, the filtered estimates of all the speakers can add up to the mixture to satisfy the above constraint. We show that this loss can promote unsupervised separation of speakers. The linear filters are computed in each sub-band based on the mixture and DNN estimates through the forward convolutive prediction (FCP) algorithm. To address the frequency permutation problem incurred by using sub-band FCP, a loss term based on minimizing intra-source magnitude scattering is proposed. Although UNSSOR requires over-determined training mixtures, we can train DNNs to achieve under-determined separation (e.g., unsupervised monaural speech separation). Evaluation results on two-speaker separation in reverberant conditions show the effectiveness and potential of UNSSOR. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ddf1dd5-cd06-4343-b220-694088717d74Cited by top-tier papers1
Ask how each one uses itBuilds on6
- PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkDacheng Yin, Chong Luo, Zhiwei Xiong, Wenjun ZengAAAI 2020 · 387 citations
- Unsupervised Sound Separation Using Mixture Invariant TrainingScott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss et al.NeurIPS 2020 · 227 citations
- Voice Separation with an Unknown Number of Multiple SpeakersEliya Nachmani, Yossi Adi, Lior WolfICML 2020 · 186 citations
- Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen SoundsEfthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey et al.ICLR 2021 · 83 citations
- The Cone of Silence: Speech Separation by LocalizationTeerapat Jenrungrot, Vivek Jayaram, Steven M. Seitz, Ira Kemelmacher-ShlizermanNeurIPS 2020 · 70 citations
Related papers
- Deep Audio Priors Emerge From Harmonic Convolutional NetworksZhoutong Zhang, Yunyun Wang, Chuang Gan, Jiajun Wu et al.ICLR 2020 · 32 citations
- PGSS: Pitch-Guided Speech SeparationXiang Li, Yiwen Wang, Yifan Sun, Xihong Wu et al.AAAI 2023 · 4 citations
- SFSRNet: Super-resolution for Single-Channel Audio Source SeparationJoel Rixen, Matthias RenzAAAI 2022 · 32 citations
- SURF: Separation via Unsupervised Remixing FlowHenry Li, Robin Scheibler, Efthymios Tzinis, Matt Shannon et al.ICML 2026
- Separate in Latent Space: Unsupervised Single Image Layer SeparationYunfei Liu, Feng LuAAAI 2020 · 16 citations
