FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio Generation
Huadai Liu, Jialei Wang, Rongjie Huang, Yang Liu, Heng Lu, Zhou Zhao, Wei Xue
Abstract
Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-toaudio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While recent methods utilizing consistency-based distillation aim to achieve few-step or single-step inference, their one-step performance is constrained by curved trajectories, preventing them from surpassing traditional diffusion models. In this work, we introduce FlashAudio with rectified flows to learn straight flow for fast simulation. To alleviate the inefficient timesteps allocation and suboptimal distribution of noise, FlashAudio optimizes the time distribution of rectified flow with Bifocal Samplers and proposes immiscible flow to minimize the total distance of data-noise pairs in a batch vias assignment. Furthermore, to address the amplified accumulation error caused by the classifierfree guidance (CFG), we propose Anchored Optimization, which refines the guidance scale by anchoring it to a reference trajectory. Experimental results on text-to-audio generation demonstrate that FlashAudio's one-step generation performance surpasses the diffusion-based models with hundreds of sampling steps on audio quality and enables a sampling speed of 400x faster than real-time on a single NVIDIA 4090Ti GPU. Code will be available at https: //github.com/liuhuadai/FlashAudio . 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12125590-0c6e-49f9-814d-5fb0f277bc5bCited by top-tier papers2
- ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and EditingHuadai Liu, Kaicheng Luo, Jialei Wang, Wen Wang et al.NeurIPS 2025 · 5 citations
- VinTAGe: Joint Video and Text Conditioning for Holistic Audio GenerationSaksham Singh Kushwaha, Yapeng TianCVPR 2025
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- AudioLCM: Efficient and High-Quality Text-to-Audio Generation with Minimal Inference StepsHuadai Liu, Rongjie Huang, Yang Liu, Hengyuan Cao et al.ACM MM 2024 · 4 citations
- InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image GenerationXingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng et al.ICLR 2024 · 358 citations
- LogCD: Local-to-global Consistency Distillation for Few-step Image GenerationQingsong Xie, Zhenyi Liao, Chen Chen, Zhijie Deng et al.CVPR 2026
- MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean FlowsXiquan Li, Junxi Liu, Yuzhe Liang, Zhikang Niu et al.ACL 2026 · 25 citations
- Invertible Consistency Distillation for Text-Guided Image Editing in Around 7 StepsNikita Starodubcev, Mikhail Khoroshikh, Artem Babenko, Dmitry BaranchukNeurIPS 2024 · 18 citations
