FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio Generation
Huadai Liu, Jialei Wang, Rongjie Huang, Yang Liu, Heng Lu, Zhou Zhao, Wei Xue
摘要
Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-toaudio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While recent methods utilizing consistency-based distillation aim to achieve few-step or single-step inference, their one-step performance is constrained by curved trajectories, preventing them from surpassing traditional diffusion models. In this work, we introduce FlashAudio with rectified flows to learn straight flow for fast simulation. To alleviate the inefficient timesteps allocation and suboptimal distribution of noise, FlashAudio optimizes the time distribution of rectified flow with Bifocal Samplers and proposes immiscible flow to minimize the total distance of data-noise pairs in a batch vias assignment. Furthermore, to address the amplified accumulation error caused by the classifierfree guidance (CFG), we propose Anchored Optimization, which refines the guidance scale by anchoring it to a reference trajectory. Experimental results on text-to-audio generation demonstrate that FlashAudio's one-step generation performance surpasses the diffusion-based models with hundreds of sampling steps on audio quality and enables a sampling speed of 400x faster than real-time on a single NVIDIA 4090Ti GPU. Code will be available at https: //github.com/liuhuadai/FlashAudio . 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and EditingHuadai Liu, Kaicheng Luo, Jialei Wang, Wen Wang 等NeurIPS 2025 · 被引用 5 次
- VinTAGe: Joint Video and Text Conditioning for Holistic Audio GenerationSaksham Singh Kushwaha, Yapeng TianCVPR 2025
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- AudioLCM: Efficient and High-Quality Text-to-Audio Generation with Minimal Inference StepsHuadai Liu, Rongjie Huang, Yang Liu, Hengyuan Cao 等ACM MM 2024 · 被引用 4 次
- InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image GenerationXingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng 等ICLR 2024 · 被引用 358 次
- LogCD: Local-to-global Consistency Distillation for Few-step Image GenerationQingsong Xie, Zhenyi Liao, Chen Chen, Zhijie Deng 等CVPR 2026
- MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean FlowsXiquan Li, Junxi Liu, Yuzhe Liang, Zhikang Niu 等ACL 2026 · 被引用 25 次
- Invertible Consistency Distillation for Text-Guided Image Editing in Around 7 StepsNikita Starodubcev, Mikhail Khoroshikh, Artem Babenko, Dmitry BaranchukNeurIPS 2024 · 被引用 18 次
