RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction
Peng Liu, Dongyang Dai, Zhiyong Wu
摘要
Recent advancements in generative modeling have significantly enhanced the reconstruction of audio waveforms from various representations. While diffusion models are adept at this task, they are hindered by latency issues due to their operation at the individual sample point level and the need for numerous sampling steps. In this study, we introduce RFWave, a cutting-edge multi-band Rectified Flow approach designed to reconstruct high-fidelity audio waveforms from Mel-spectrograms or discrete acoustic tokens. RFWave uniquely generates complex spectrograms and operates at the frame level, processing all subbands simultaneously to boost efficiency. Leveraging Rectified Flow, which targets a straight transport trajectory, RFWave achieves reconstruction with just 10 sampling steps. Our empirical evaluations show that RFWave not only provides outstanding reconstruction quality but also offers vastly superior computational efficiency, enabling audio generation at speeds up to 160 times faster than real-time on a GPU. Both an online demonstration and the source code are accessible 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Flow Straight and Fast in Hilbert Space: Functional Rectified FlowJianxin Zhang, Clayton ScottICLR 2026 · 被引用 6 次
- Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio GenerationZengwei Yao, Wei Kang, Han Zhu, Liyong Guo 等ICLR 2026 · 被引用 5 次
- StreamFlow: Streaming Audio Generation from Discrete Tokens via Streaming Flow MatchingHa-Yeong Choi, Sang-Hoon LeeNeurIPS 2025 · 被引用 2 次
- BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching ModelsSusan Liang, Dejan Markovic, Israel D. Gebru, Steven Krenn 等ICML 2025
- Toward Complex-Valued Neural Networks for Waveform GenerationHyung-Seok Oh, Deok-Hyeon Cho, Seung-Bin Kim, Seong-Whan LeeICLR 2026
它引用的顶会 Paper14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Neural Networks Fail to Learn Periodic Functions and How to Fix ItLiu Ziyin, Tilman Hartwig, Masahito UedaNeurIPS 2020 · 被引用 249 次
相关 Paper
- WaveFlow: A Compact Flow-based Model for Raw AudioWei Ping, Kainan Peng, Kexin Zhao, Zhao SongICML 2020 · 被引用 132 次
- FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio GenerationHuadai Liu, Jialei Wang, Rongjie Huang, Yang Liu 等ACL 2025 · 被引用 16 次
- ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion ModelsSang-Hoon Lee, Ha-Yeong ChoiICML 2026
- Wavelet Diffusion Models are fast and scalable Image GeneratorsHao Phung, Quan Dao, Anh TranCVPR 2023
- PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform GenerationSang-Hoon Lee, Ha-Yeong Choi, Seong-Whan LeeICLR 2025
