AudioLCM: Efficient and High-Quality Text-to-Audio Generation with Minimal Inference Steps
Huadai Liu, Rongjie Huang, Yang Liu, Hengyuan Cao, Jialei Wang, Xize Cheng, Siqi Zheng, Zhou Zhao
摘要
Recent advancements in Latent Diffusion Models (LDMs) have propelled them to the forefront of various generative tasks. However, their iterative sampling process poses a significant computational burden, resulting in slow generation speeds and limiting their application in text-to-audio generation deployment. In this work, we introduce AudioLCM, a novel consistency-based model tailored for efficient and high-quality text-to-audio generation. Unlike prior approaches that address noise removal through iterative processes, AudioLCM integrates Consistency Models (CMs) into the generation process, facilitating rapid inference through a mapping from any point at any time step to the trajectory's initial point. To overcome the convergence issue inherent in LDMs with reduced sample iterations, we propose the Guided Latent Consistency Distillation with a multi-step Ordinary Differential Equation (ODE) solver. This innovation shortens the time schedule from thousands to dozens of steps while maintaining sample quality, thereby achieving fast convergence and high-quality generation. Furthermore, to optimize the performance of transformer-based neural network architectures, we integrate the advanced techniques pioneered by LLaMA into the foundational framework of transformers. This architecture supports stable and efficient training, ensuring robust performance in text-to-audio synthesis. Experimental results on text-to-audio generation and text-to-music synthesis tasks demonstrate that AudioLCM needs only 2 iterations to synthesize high-fidelity audios, while it maintains sample quality competitive with state-of-the-art models using hundreds of steps. AudioLCM enables a sampling speed of 333x faster than real-time on a single NVIDIA 4090Ti GPU, making generative models practically applicable to text-to-audio generation deployment. Our extensive preliminary analysis shows that each design in AudioLCM is effective. https://AudioLCM.github.io/. Code is Available https://github.com/Text-to-Audio/AudioLCM
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and EditingHuadai Liu, Kaicheng Luo, Jialei Wang, Wen Wang 等NeurIPS 2025 · 被引用 5 次
- SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for VideoSuhyeon Yoo, Adolfo Hernandez Santisteban, Prem Seetharaman, Justin Salamon 等CHI 2026 · 被引用 2 次
- OmniAudio: Generating Spatial Audio from 360-Degree VideoHuadai Liu, Tianyi Luo, Kaicheng Luo, Qikai Jiang 等ICML 2025
相关 Paper
- FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio GenerationHuadai Liu, Jialei Wang, Rongjie Huang, Yang Liu 等ACL 2025 · 被引用 16 次
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
- Continuous Audio Language ModelsSimon Rouard, Manu Orsini, Axel Roebel, Neil Zeghidour 等ICLR 2026 · 被引用 13 次
- IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion ModelingKuan-Po Huang, Shu-Wen Yang, Huy Phan, Bo-Ru Lu 等ICML 2025
- ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-SpeechRongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu 等ACM MM 2022 · 被引用 182 次
