Separate and Diffuse: Using a Pretrained Diffusion Model for Better Source Separation
Shahar Lutati, Eliya Nachmani, Lior Wolf
Abstract
The problem of speech separation, also known as the cocktail party problem, refers to the task of isolating a single speech signal from a mixture of speech signals. Previous work on source separation derived an upper bound for the source separation task in the domain of human speech. This bound is derived for deterministic models. Recent advancements in generative models challenge this bound. We show how the upper bound can be generalized to the case of random generative models. Applying a diffusion model Vocoder that was pretrained to model single-speaker voices on the output of a deterministic separation model leads to state-of-the-art separation results. It is shown that this requires one to combine the output of the separation model with that of the diffusion model. In our method, a linear combination is performed, in the frequency domain, using weights that are inferred by a learned model. We show state-of-the-art results on 2, 3, 5, 10, and 20 speakers on multiple benchmarks. In particular, for two speakers, our method is able to surpass what was previously considered the upper performance bound.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4115dd6e-6d49-4282-8100-b8a57e56b306Cited by top-tier papers2
- Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech SeparationUi-Hyeop Shin, Sangyoun Lee, Taehan Kim, Hyung-Min ParkNeurIPS 2024 · 46 citations
- ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion PriorZhongweiyang Xu, Xulin Fan, Zhong-Qiu Wang, Xilin Jiang et al.ICML 2025
Builds on6
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- High Fidelity Speech Synthesis with Adversarial NetworksMikolaj Binkowski, Jeff Donahue, Sander Dieleman, Aidan Clark et al.ICLR 2020 · 263 citations
- Voice Separation with an Unknown Number of Multiple SpeakersEliya Nachmani, Yossi Adi, Lior WolfICML 2020 · 186 citations
- Source Separation with Deep Generative PriorsVivek Jayaram, John ThickstunICML 2020 · 47 citations
Related papers
- ZeroSep: Separate Anything in Audio with Zero TrainingChao Huang, Yuesheng Ma, Junxuan Huang, Susan Liang et al.NeurIPS 2025 · 8 citations
- Multi-Source Diffusion Models for Simultaneous Music Generation and SeparationGiorgio Mariani, Irene Tallini, Emilian Postolache, Michele Mancusi et al.ICLR 2024 · 75 citations
- Unsupervised Single-Channel Audio Separation with Diffusion Source PriorsRunwu Shi, Chang Li, Jiang Wang, Rui Zhang et al.AAAI 2026
- A Mixture-Based Framework for Guiding Diffusion ModelsYazid Janati, Badr Moufad, Mehdi Abou El Qassime, Alain Oliviero Durmus et al.ICML 2025
- Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling SchemeVadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova et al.ICLR 2022 · 185 citations
