Complex-Cycle-Consistent Diffusion Model for Monaural Speech Enhancement
Yi Li, Yang Sun, Plamen P. Angelov
Abstract
In this paper, we present a novel diffusion model-based monaural speech enhancement method. Our approach incorporates the separate estimation of speech spectra's magnitude and phase in two diffusion networks. Throughout the diffusion process, noise clips from real-world noise interferences are added gradually to the clean speech spectra and a noise-aware reverse process is proposed to learn how to generate both clean speech spectra and noise spectra. Furthermore, to fully leverage the intrinsic relationship between magnitude and phase, we introduce a complex-cycle-consistent (CCC) mechanism that uses the estimated magnitude to map the phase, and vice versa. We implement this algorithm within a phase-aware speech enhancement diffusion model (SEDM). We conduct extensive experiments on public datasets to demonstrate the effectiveness of our method, highlighting the significant benefits of exploiting the intrinsic relationship between phase and magnitude information to enhance speech. The comparison to conventional diffusion models demonstrates the superiority of SEDM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac1aeea4-2376-4374-9c98-5483e11abe1bCited by top-tier papers1
Ask how each one uses itBuilds on3
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Realistic Noise Synthesis with Diffusion ModelsQi Wu, Mingyan Han, Ting Jiang, Chengzhi Jiang et al.AAAI 2025 · 6 citations
- Ambiguous Medical Image Segmentation Using Diffusion ModelsAimon Rahman, Jeya Maria Jose Valanarasu, Ilker Hacihaliloglu, Vishal M. PatelCVPR 2023
Related papers
- PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkDacheng Yin, Chong Luo, Zhiwei Xiong, Wenjun ZengAAAI 2020 · 387 citations
- BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech EnhancementCunhang Fan, Enrui Liu, Andong Li, Jianhua Tao et al.AAAI 2025
- Frequency Domain-Based Diffusion Model for Unpaired Image DehazingChengxu Liu, Lu Qi, Jinshan Pan, Xueming Qian et al.ICCV 2025 · 13 citations
- Interactive Speech and Noise Modeling for Speech EnhancementChengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan et al.AAAI 2021 · 112 citations
- ClearSpeech: Improving Voice Quality of Earbuds Using Both In-Ear and Out-Ear MicrophonesDong Ma, Ting Dang, Ming Ding, Rajesh BalanUbiComp 2024 · 5 citations
