Text-to-Audio Generation using Instruction Guided Latent Diffusion Model
Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, Soujanya Poria
Abstract
The immense scale of the recent large language models (LLM) allows many interesting properties, such as, instruction- and chain-of-thought-based fine-tuning, that has significantly improved zero- and few-shot performance in many natural language processing (NLP) tasks. Inspired by such successes, we adopt such an instruction-tuned LLM Flan-T5 as the text encoder for text-to-audio (TTA) generation-a task where the goal is to generate an audio from its textual description. The prior works on TTA either pre-trained a joint text-audio encoder or used a non-instruction-tuned model, such as, T5. Consequently, our latent diffusion model (LDM)-based approach (Tango) outperforms the state-of-the-art AudioLDM on most metrics and stays comparable on the rest on AudioCaps test set, despite training the LDM on a 63 times smaller dataset and keeping the text encoder frozen. This improvement might also be attributed to the adoption of audio pressure level-based sound mixing for training set augmentation, whereas the prior methods take a random mix.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get af532be2-06d6-4e62-8b03-15b3417a1178Cited by top-tier papers19
- V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation ModelsHeng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright et al.AAAI 2024 · 84 citations
- Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM InversionHila Manor, Tomer MichaeliICML 2024 · 61 citations
- MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean FlowsXiquan Li, Junxi Liu, Yuzhe Liang, Zhikang Niu et al.ACL 2026 · 25 citations
- "AI Afterlife" as Digital Legacy: Perceptions, Expectations, and ConcernsYing Lei, Shuai Ma, Yuling Sun, Xiaojuan MaCHI 2025 · 24 citations
- Prompt-guided Precise Audio Editing with Diffusion ModelsManjie Xu, Chenxing Li, Duzhen Zhang, Dan Su et al.ICML 2024 · 15 citations
Related papers
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
- Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token PredictionShu-Wen Yang, Byeonggeun Kim, Kuan-Po Huang, Qingming Tang et al.ICML 2025
- Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic DataSreyan Ghosh, Sonal Kumar, Zhifeng Kong, Rafael Valle et al.ICLR 2025
- Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference OptimizationNavonil Majumder, Chia-Yu Hung, Deepanway Ghosal, Wei-Ning Hsu et al.ACM MM 2024 · 52 citations
- IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion ModelingKuan-Po Huang, Shu-Wen Yang, Huy Phan, Bo-Ru Lu et al.ICML 2025
