Rethinking Flow and Diffusion Bridge Models for Speech Enhancement
Dahan Wang, Jun Gao, Tong Lei, Yuxiang Hu, Changbao Zhu, Kai Chen, Jing Lu
摘要
Flow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech signals based on principles such as flow matching, score matching, and Schrödinger bridge. In this paper, we present a framework that unifies existing flow and diffusion bridge models by interpreting them as constructions of Gaussian probability paths with varying means and variances between paired data. Furthermore, we investigate the underlying consistency between the training/inference procedures of these generative models and conventional predictive models. Our analysis reveals that each sampling step of a well-trained flow or diffusion bridge model optimized with a data prediction loss is theoretically analogous to executing predictive speech enhancement. Motivated by this insight, we introduce an enhanced bridge model that integrates an effective probability path design with key elements from predictive paradigms, including improved network architecture, tailored loss functions, and optimized training strategies. Experiments on denoising and dereverberation tasks demonstrate that the proposed method outperforms existing flow and diffusion baselines with fewer parameters and reduced computational complexity. The results also highlight that the inherently predictive nature of this generative framework imposes limitations on its achievable upper-bound performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkDacheng Yin, Chong Luo, Zhiwei Xiong, Wenjun ZengAAAI 2020 · 被引用 387 次
- Interactive Speech and Noise Modeling for Speech EnhancementChengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan 等AAAI 2021 · 被引用 112 次
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel 等ICLR 2023 · 被引用 87 次
- Score-based Generative Modeling through Stochastic Evolution Equations in Hilbert SpacesSungbin Lim, Eun-Bi Yoon, Taehyun Byun, Taewon Kang 等NeurIPS 2023 · 被引用 55 次
相关 Paper
- Denoising Diffusion Bridge ModelsLinqi Zhou, Aaron Lou, Samar Khanna, Stefano ErmonICLR 2024 · 被引用 163 次
- Denoising Diffusion SamplersFrancisco Vargas, Will Sussman Grathwohl, Arnaud DoucetICLR 2023 · 被引用 3 次
- Revisiting Denoising Diffusion Probabilistic Models for Speech Enhancement: Condition Collapse, Efficiency and RefinementWenxin Tai, Fan Zhou, Goce Trajcevski, Ting ZhongAAAI 2023 · 被引用 38 次
- System-Embedded Diffusion Bridge ModelsBartlomiej Sobieski, Matthew Tivnan, Yuang Wang, Siyeop Yoon 等NeurIPS 2025 · 被引用 4 次
- Consistency Diffusion Bridge ModelsGuande He, Kaiwen Zheng, Jianfei Chen, Fan Bao 等NeurIPS 2024 · 被引用 31 次
