Cauchy Diffusion: A Heavy-tailed Denoising Diffusion Probabilistic Model for Speech Synthesis
Qi Lian, Yu Qi, Yueming Wang
Abstract
Denoising diffusion probabilistic models (DDPMs) have gained popularity in devising neural vocoders and obtained outstanding performance. However, existing DDPM-based neural vocoders struggle to handle the prosody diversities due to their susceptibility to mode-collapse issues confronted with imbalanced data. We introduced Cauchy Diffusion, a model incorporating the Cauchy noises to address this challenge. The heavy-tailed Cauchy distribution exhibits better resilience to imbalanced speech data, potentially improving prosody modeling. Our experiments on the LJSpeech and VCTK datasets demonstrate that Cauchy Diffusion achieved state-of-the-art speech synthesis performance. Compared to existing neural vocoders, our Cauchy Diffusion notably improved speech diversity while maintaining superior speech quality. Remarkably, Cauchy Diffusion surpassed neural vocoders based on generative adversarial networks (GANs) that are explicitly optimized to improve diversity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72444635-7b54-4b64-9d9a-667db4d437ecCited by top-tier papers4
- Listening Through the Noise: Cauchy-Driven Diffusion Bridges for Robust Gastrointestinal Auscultation and Clinical BenchmarkingDian Ding, Liren Dong, Yu Lu, Juntao Zhou et al.ICML 2026
- L-Diffusion: Laplace Diffusion for Efficient Pathology Image SegmentationWeihan Li, Linyun Zhou, Yang Jian, Shengxuming Zhang et al.ICML 2025
- Condition-Aware Graph Flow Matching for Modeling the Distributions of Complex Fluid SystemsXiaochao Deng, Jie Chen, Xiaogang DengICML 2026
- Pareto Variational AutoencoderMincheol Cho, Yedarm Seong, Joong-Ho WonICLR 2026
Builds on13
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- Scalable Adaptive Computation for Iterative GenerationAllan Jabri, David J. Fleet, Ting ChenICML 2023 · 175 citations
Related papers
- Revisiting Denoising Diffusion Probabilistic Models for Speech Enhancement: Condition Collapse, Efficiency and RefinementWenxin Tai, Fan Zhou, Goce Trajcevski, Ting ZhongAAAI 2023 · 38 citations
- PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive PriorSang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan et al.ICLR 2022 · 117 citations
- DPP-TTS: Diversifying prosodic features of speech via determinantal point processesSeongho Joo, Hyukhun Koh, Kyomin JungEMNLP 2023
- Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion ModelXiangyu Zhang, Daijiao Liu, Hexin Liu, Qiquan Zhang et al.EMNLP 2024 · 3 citations
- Heavy-Tailed Diffusion with Denoising Levy Probabilistic ModelsDario Shariatian, Umut Simsekli, Alain Oliviero DurmusICLR 2025
