ZeroSep: Separate Anything in Audio with Zero Training
Chao Huang, Yuesheng Ma, Junxuan Huang, Susan Liang, Yunlong Tang, Jing Bi, Wenqiang Liu, Nima Mesgarani, Chenliang Xu
摘要
Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive, task-specific labeled data and struggle to generalize to the immense variability and open-set nature of real-world acoustic scenes. Inspired by the success of generative foundation models, we investigate whether pre-trained text-guided audio diffusion models can overcome these limitations. We make a surprising discovery: zero-shot source separation can be achieved purely through a pre-trained text-guided audio diffusion model under the right configuration. Our method, named ZeroSep, works by inverting the mixed audio into the diffusion model's latent space and then using text conditioning to guide the denoising process to recover individual sources. Without any task-specific training or fine-tuning, ZeroSep repurposes the generative diffusion model for a discriminative separation task and inherently supports openset scenarios through its rich textual priors. ZeroSep is compatible with a variety of pre-trained text-guided audio diffusion backbones and delivers strong separation performance on multiple separation benchmarks, surpassing even supervised methods. Our project page is here: https://wikichao.github.io/ZeroSep/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
相关 Paper
- Unsupervised Single-Channel Audio Separation with Diffusion Source PriorsRunwu Shi, Chang Li, Jiang Wang, Rui Zhang 等AAAI 2026
- Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataKe Chen, Xingjian Du, Bilei Zhu, Zejun Ma 等AAAI 2022 · 被引用 58 次
- OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio SeparationTanvir Mahmud, Diana MarculescuEMNLP 2024 · 被引用 1 次
- A Mixture-Based Framework for Guiding Diffusion ModelsYazid Janati, Badr Moufad, Mehdi Abou El Qassime, Alain Oliviero Durmus 等ICML 2025
- SAM Audio: Segment Anything in AudioBowen Shi, Andros Tjandra, John Hoffman, Helin Wang 等ICML 2026 · 被引用 35 次
