Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model
Xiangyu Zhang, Daijiao Liu, Hexin Liu, Qiquan Zhang, Hanyu Meng, Leibny Paola García-Perera, Engsiong Chng, Lina Yao
摘要
Recently, Denoising Diffusion Probabilistic Models (DDPMs) have attained leading performances across a diverse range of generative tasks. However, in the field of speech synthesis, although DDPMs exhibit impressive performance, their long training duration and substantial inference costs hinder practical deployment. Existing approaches primarily focus on enhancing inference speed, while approaches to accelerate training-a key factor in the costs associated with adding or customizing voices-often necessitate complex modifications to the model, compromising their universal applicability. To address the aforementioned challenges, we propose an inquiry: is it possible to enhance the training/inference speed and performance of DDPMs by modifying the speech signal itself? In this paper, we double the training and inference speed of Speech DDPMs by simply redirecting the generative target to the wavelet domain. This method not only achieves comparable or superior performance to the original model in speech synthesis tasks but also demonstrates its versatility. By investigating and utilizing different wavelet bases, our approach proves effective not just in speech synthesis, but also in speech enhancement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music EditingXinlei Niu, Kin Wai Cheuk, Jing Zhang, Naoki Murata 等AAAI 2026 · 被引用 5 次
- GenSE: Generative Speech Enhancement via Language Models using Hierarchical ModelingJixun Yao, Hexin Liu, Chen Chen, Yuchen Hu 等ICLR 2025
- Wavelet Predictive Representations for Non-Stationary Reinforcement LearningMin Wang, Xin Li, Ye He, Yao-Hui Li 等ICLR 2026
它引用的顶会 Paper16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Revisiting Denoising Diffusion Probabilistic Models for Speech Enhancement: Condition Collapse, Efficiency and RefinementWenxin Tai, Fan Zhou, Goce Trajcevski, Ting ZhongAAAI 2023 · 被引用 38 次
- WaveEx: Accelerating Flow Matching-based Speech Generation via Wavelet-guided ExtrapolationXiaoqian Liu, Xiyan Gui, Zhengkun Ge, Yuan Ge 等AAAI 2026 · 被引用 1 次
- PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive PriorSang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan 等ICLR 2022 · 被引用 117 次
- CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency ModelZhen Ye, Wei Xue, Xu Tan, Jie Chen 等ACM MM 2023 · 被引用 30 次
- Wavelet Diffusion Models are fast and scalable Image GeneratorsHao Phung, Quan Dao, Anh TranCVPR 2023
