WaveFlow: A Compact Flow-based Model for Raw Audio
Wei Ping, Kainan Peng, Kexin Zhao, Zhao Song
摘要
In this work, we propose WaveFlow, a small-footprint generative flow for raw audio, which is directly trained with maximum likelihood. It handles the long-range structure of 1-D waveform with a dilated 2-D convolutional architecture, while modeling the local variations using expressive autoregressive functions. WaveFlow provides a unified view of likelihood-based models for 1-D data, including WaveNet and WaveGlow as special cases. It generates high-fidelity speech as WaveNet, while synthesizing several orders of magnitude faster as it only requires a few sequential steps to generate very long waveforms with hundreds of thousands of time-steps. Furthermore, it can significantly reduce the likelihood gap that has existed between autoregressive models and flow-based models for efficient synthesis. Finally, our small-footprint WaveFlow has only 5.91M parameters, which is 15 smaller than WaveGlow. It can generate 22.05 kHz high-fidelity audio 42.6 faster than real-time (at a rate of 939.3 kHz) on a V100 GPU without engineered inference kernels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- It's Raw! Audio Generation with State-Space ModelsKaran Goel, Albert Gu, Chris Donahue, Christopher RéICML 2022 · 被引用 257 次
- Likelihood Training of Schrödinger Bridge using Forward-Backward SDEs TheoryTianrong Chen, Guan-Horng Liu, Evangelos A. TheodorouICLR 2022 · 被引用 249 次
- D2C: Diffusion-Decoding Models for Few-Shot Conditional GenerationAbhishek Sinha, Jiaming Song, Chenlin Meng, Stefano ErmonNeurIPS 2021 · 被引用 149 次
- VAEBM: A Symbiosis between Variational Autoencoders and Energy-based ModelsZhisheng Xiao, Karsten Kreis, Jan Kautz, Arash VahdatICLR 2021 · 被引用 139 次
它引用的顶会 Paper1
相关 Paper
- DFlow: A Generative Model Combining Denoising AutoEncoder and Normalizing Flow for High Fidelity Waveform GenerationChenfeng Miao, Qingying Zhu, Minchuan Chen, Wei Hu 等ICML 2024 · 被引用 2 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- WaveGrad: Estimating Gradients for Waveform GenerationNanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss 等ICLR 2021 · 被引用 44 次
- RFWave: Multi-band Rectified Flow for Audio Waveform ReconstructionPeng Liu, Dongyang Dai, Zhiyong WuICLR 2025
- ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion ModelsSang-Hoon Lee, Ha-Yeong ChoiICML 2026
