Cauchy Diffusion: A Heavy-tailed Denoising Diffusion Probabilistic Model for Speech Synthesis
Qi Lian, Yu Qi, Yueming Wang
摘要
Denoising diffusion probabilistic models (DDPMs) have gained popularity in devising neural vocoders and obtained outstanding performance. However, existing DDPM-based neural vocoders struggle to handle the prosody diversities due to their susceptibility to mode-collapse issues confronted with imbalanced data. We introduced Cauchy Diffusion, a model incorporating the Cauchy noises to address this challenge. The heavy-tailed Cauchy distribution exhibits better resilience to imbalanced speech data, potentially improving prosody modeling. Our experiments on the LJSpeech and VCTK datasets demonstrate that Cauchy Diffusion achieved state-of-the-art speech synthesis performance. Compared to existing neural vocoders, our Cauchy Diffusion notably improved speech diversity while maintaining superior speech quality. Remarkably, Cauchy Diffusion surpassed neural vocoders based on generative adversarial networks (GANs) that are explicitly optimized to improve diversity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Listening Through the Noise: Cauchy-Driven Diffusion Bridges for Robust Gastrointestinal Auscultation and Clinical BenchmarkingDian Ding, Liren Dong, Yu Lu, Juntao Zhou 等ICML 2026
- L-Diffusion: Laplace Diffusion for Efficient Pathology Image SegmentationWeihan Li, Linyun Zhou, Yang Jian, Shengxuming Zhang 等ICML 2025
- Condition-Aware Graph Flow Matching for Modeling the Distributions of Complex Fluid SystemsXiaochao Deng, Jie Chen, Xiaogang DengICML 2026
- Pareto Variational AutoencoderMincheol Cho, Yedarm Seong, Joong-Ho WonICLR 2026
它引用的顶会 Paper13
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- Scalable Adaptive Computation for Iterative GenerationAllan Jabri, David J. Fleet, Ting ChenICML 2023 · 被引用 175 次
相关 Paper
- Revisiting Denoising Diffusion Probabilistic Models for Speech Enhancement: Condition Collapse, Efficiency and RefinementWenxin Tai, Fan Zhou, Goce Trajcevski, Ting ZhongAAAI 2023 · 被引用 38 次
- PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive PriorSang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan 等ICLR 2022 · 被引用 117 次
- DPP-TTS: Diversifying prosodic features of speech via determinantal point processesSeongho Joo, Hyukhun Koh, Kyomin JungEMNLP 2023
- Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion ModelXiangyu Zhang, Daijiao Liu, Hexin Liu, Qiquan Zhang 等EMNLP 2024 · 被引用 3 次
- Heavy-Tailed Diffusion with Denoising Levy Probabilistic ModelsDario Shariatian, Umut Simsekli, Alain Oliviero DurmusICLR 2025
