RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction
Peng Liu, Dongyang Dai, Zhiyong Wu
Abstract
Recent advancements in generative modeling have significantly enhanced the reconstruction of audio waveforms from various representations. While diffusion models are adept at this task, they are hindered by latency issues due to their operation at the individual sample point level and the need for numerous sampling steps. In this study, we introduce RFWave, a cutting-edge multi-band Rectified Flow approach designed to reconstruct high-fidelity audio waveforms from Mel-spectrograms or discrete acoustic tokens. RFWave uniquely generates complex spectrograms and operates at the frame level, processing all subbands simultaneously to boost efficiency. Leveraging Rectified Flow, which targets a straight transport trajectory, RFWave achieves reconstruction with just 10 sampling steps. Our empirical evaluations show that RFWave not only provides outstanding reconstruction quality but also offers vastly superior computational efficiency, enabling audio generation at speeds up to 160 times faster than real-time on a GPU. Both an online demonstration and the source code are accessible 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a47219ad-0d14-4959-bba2-1e945b53f75cCited by top-tier papers8
- Flow Straight and Fast in Hilbert Space: Functional Rectified FlowJianxin Zhang, Clayton ScottICLR 2026 · 6 citations
- Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio GenerationZengwei Yao, Wei Kang, Han Zhu, Liyong Guo et al.ICLR 2026 · 5 citations
- StreamFlow: Streaming Audio Generation from Discrete Tokens via Streaming Flow MatchingHa-Yeong Choi, Sang-Hoon LeeNeurIPS 2025 · 2 citations
- BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching ModelsSusan Liang, Dejan Markovic, Israel D. Gebru, Steven Krenn et al.ICML 2025
- Toward Complex-Valued Neural Networks for Waveform GenerationHyung-Seok Oh, Deok-Hyeon Cho, Seung-Bin Kim, Seong-Whan LeeICLR 2026
Builds on14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Neural Networks Fail to Learn Periodic Functions and How to Fix ItLiu Ziyin, Tilman Hartwig, Masahito UedaNeurIPS 2020 · 249 citations
Related papers
- WaveFlow: A Compact Flow-based Model for Raw AudioWei Ping, Kainan Peng, Kexin Zhao, Zhao SongICML 2020 · 132 citations
- FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio GenerationHuadai Liu, Jialei Wang, Rongjie Huang, Yang Liu et al.ACL 2025 · 16 citations
- ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion ModelsSang-Hoon Lee, Ha-Yeong ChoiICML 2026
- Wavelet Diffusion Models are fast and scalable Image GeneratorsHao Phung, Quan Dao, Anh TranCVPR 2023
- PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform GenerationSang-Hoon Lee, Ha-Yeong Choi, Seong-Whan LeeICLR 2025
