DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
Roi Benita, Michael Elad, Joseph Keshet
摘要
Diffusion models have recently been shown to be relevant for high-quality speech generation. Most work has been focused on generating spectrograms, and as such, they further require a subsequent model to convert the spectrogram to a waveform (i.e., a vocoder). This work proposes a diffusion probabilistic end-to-end model for directly generating the raw speech waveform. The proposed model is autoregressive, generating overlapping frames sequentially, where each frame is conditioned on a portion of the previously generated one. Hence, our model can effectively synthesize an unlimited speech duration while preserving high-fidelity synthesis and temporal coherence. We implemented the proposed model for unconditional and conditional speech generation, where the latter can be driven by an input sequence of phonemes, amplitudes, and pitch values. Working directly on the waveform has some empirical advantages. Specifically, it allows the creation of local acoustic behaviors, like vocal fry, which makes the overall waveform sounds more natural. Furthermore, the proposed diffusion model is stochastic and not deterministic; therefore, each inference generates a slightly different waveform variation, enabling abundance of valid realizations. Experiments show that the proposed model generates speech with superior quality compared with other state-of-the-art neural speech generation systems. 1 2
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion RefinementChenyu Yang, Shuai Wang, Hangting Chen, Wei Tan 等NeurIPS 2025 · 被引用 28 次
- Fine-grained Control of Generative Data Augmentation in IoT SensingTianshi Wang, Qikai Yang, Ruijie Wang, Dachun Sun 等NeurIPS 2024 · 被引用 13 次
- Spectral Analysis of Diffusion Models with Application to Schedule DesignRoi Benita, Miki Elad, Joseph KeshetNeurIPS 2025 · 被引用 13 次
它引用的顶会 Paper17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- WaveGrad: Estimating Gradients for Waveform GenerationNanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss 等ICLR 2021 · 被引用 44 次
- DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech TranslationYongxin Zhu, Zhujin Gao, Xinyuan Zhou, Zhongyi Ye 等EMNLP 2023 · 被引用 1 次
- End-to-end Adversarial Text-to-SpeechJeff Donahue, Sander Dieleman, Mikolaj Binkowski, Erich Elsen 等ICLR 2021 · 被引用 33 次
- PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive PriorSang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan 等ICLR 2022 · 被引用 117 次
