Revisiting Denoising Diffusion Probabilistic Models for Speech Enhancement: Condition Collapse, Efficiency and Refinement
Wenxin Tai, Fan Zhou, Goce Trajcevski, Ting Zhong
摘要
Recent literature has shown that denoising diffusion probabilistic models (DDPMs) can be used to synthesize high-fidelity samples with a competitive (or sometimes better) quality than previous state-of-the-art approaches. However, few attempts have been made to apply DDPM for the speech enhancement task. The reported performance of the existing works is relatively poor and significantly inferior to other generative methods. In this work, we first reveal the difficulties in applying existing diffusion models to the field of speech enhancement. Then we introduce DR-DiffuSE, a simple and effective framework for speech enhancement using conditional diffusion models. We present three strategies (two in diffusion training and one in reverse sampling) to tackle the condition collapse and guarantee the sufficient use of condition information. For efficiency, we introduce the fast sampling technique to reduce the sampling process into several steps and exploit a refinement network to calibrate the defective speech. Our proposed method achieves the state-of-the-art performance to the GAN-based model and shows a significant improvement over existing DDPM-based algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DOSE: Diffusion Dropout with Adaptive Prior for Speech EnhancementWenxin Tai, Yue Lei, Fan Zhou, Goce Trajcevski 等NeurIPS 2023 · 被引用 39 次
- Regularized Conditional Diffusion Model for Multi-Task Preference AlignmentXudong Yu, Chenjia Bai, Haoran He, Changhong Wang 等NeurIPS 2024 · 被引用 11 次
- Rethinking Flow and Diffusion Bridge Models for Speech EnhancementDahan Wang, Jun Gao, Tong Lei, Yuxiang Hu 等AAAI 2026 · 被引用 1 次
- RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned PriorChing Hua Lee, Chouchang Yang, Jaejin Cho, Yashas Malur Saidutta 等ICML 2025
它引用的顶会 Paper4
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- Deblurring via Stochastic RefinementJay Whang, Mauricio Delbracio, Hossein Talebi, Chitwan Saharia 等CVPR 2022
相关 Paper
- Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion ModelXiangyu Zhang, Daijiao Liu, Hexin Liu, Qiquan Zhang 等EMNLP 2024 · 被引用 3 次
- CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency ModelZhen Ye, Wei Xue, Xu Tan, Jie Chen 等ACM MM 2023 · 被引用 30 次
- PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive PriorSang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan 等ICLR 2022 · 被引用 117 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-SpeechRongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu 等ACM MM 2022 · 被引用 182 次
