Fugatto 1: Foundational Generative Audio Transformer Opus 1
Rafael Valle, Rohan Badlani, Zhifeng Kong, Sang-gil Lee, Arushi Goel, Sungwon Kim, João Felipe Santos, Shuqi Dai, Siddharth Gururani, Aya Aljafari, Alexander H. Liu, Kevin J. Shih
Abstract
Fugatto is a versatile audio synthesis and transformation model capable of following free-form text instructions with optional audio inputs. While large language models (LLMs) trained with text on a simple next-token prediction objective can learn to infer instructions directly from the data, models trained solely on audio data lack this capacity. This is because audio data does not inherently contain the instructions that were used to generate it. To overcome this challenge, we introduce a specialized dataset generation approach optimized for producing a wide range of audio generation and transformation tasks, ensuring the data reveals meaningful relationships between audio and language. Another challenge lies in achieving compositional abilities -such as combining, interpolating between, or negating instructions -using data alone. To address it, we propose ComposableART, an inference-time technique that extends classifier-free guidance to compositional guidance. It enables the seamless and flexible composition of instructions, leading to highly customizable audio outputs outside the training distribution. Our evaluations across a diverse set of tasks demonstrate that Fugatto performs competitively with specialized models, while ComposableART enhances its sonic palette and control over synthesis. Most notably, we highlight our framework's ability to synthesize emergent sounds -sonic phenomena that transcend conventional audio generation -unlocking new creative possibilities. Demo Website.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91924f61-9f31-4db2-97d2-bf8a69ebf9f8Cited by top-tier papers5
- UALM: Unified Audio Language Model for Understanding, Generation and ReasoningJinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh et al.ICLR 2026 · 17 citations
- SAO-Instruct: Free-form Audio Editing using Natural Language InstructionsMichael Ungersböck, Florian Grötschla, Luca A. Lanzendörfer, June Young Yi et al.NeurIPS 2025 · 9 citations
- AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion ForcingWilliam Chen, Prem Seetharaman, Rithesh Kumar, Oriol Nieto et al.ICML 2026 · 7 citations
- AudioStory: Generating Long-Form Narrative Audio with Large Language ModelsYuxin Guo, Teng Wang, Yuying Ge, Shijie Ma et al.CVPR 2026 · 5 citations
- SpeechOp: Inference-Time Task Composition for Generative Speech ProcessingJustin Lovelace, Rithesh Kumar, Jiaqi Su, Ke Chen et al.ICLR 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Guiding a Diffusion Model with a Bad Version of ItselfTero Karras, Miika Aittala, Tuomas Kynkäänniemi, Jaakko Lehtinen et al.NeurIPS 2024 · 338 citations
- Listen, Think, and UnderstandYuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky et al.ICLR 2024 · 247 citations
- Compositional Visual Generation with Energy Based ModelsYilun Du, Shuang Li, Igor MordatchNeurIPS 2020 · 225 citations
- Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue AbilitiesZhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping et al.ICML 2024 · 207 citations
Related papers
- Text-to-Audio Generation using Instruction Guided Latent Diffusion ModelDeepanway Ghosal, Navonil Majumder, Ambuj Mehrish, Soujanya PoriaACM MM 2023 · 90 citations
- BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMsYue Wang, Ruotian Ma, Xingyu Chen, Zhengliang Shi et al.ACL 2026
- UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text InstructionsChunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang et al.ACL 2026 · 2 citations
- Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic DataSreyan Ghosh, Sonal Kumar, Zhifeng Kong, Rafael Valle et al.ICLR 2025
- FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio GenerationYuxuan Jiang, Zehua Chen, Zeqian Ju, Chang Li et al.ACM MM 2025 · 5 citations
