Fugatto 1: Foundational Generative Audio Transformer Opus 1
Rafael Valle, Rohan Badlani, Zhifeng Kong, Sang-gil Lee, Arushi Goel, Sungwon Kim, João Felipe Santos, Shuqi Dai, Siddharth Gururani, Aya Aljafari, Alexander H. Liu, Kevin J. Shih
摘要
Fugatto is a versatile audio synthesis and transformation model capable of following free-form text instructions with optional audio inputs. While large language models (LLMs) trained with text on a simple next-token prediction objective can learn to infer instructions directly from the data, models trained solely on audio data lack this capacity. This is because audio data does not inherently contain the instructions that were used to generate it. To overcome this challenge, we introduce a specialized dataset generation approach optimized for producing a wide range of audio generation and transformation tasks, ensuring the data reveals meaningful relationships between audio and language. Another challenge lies in achieving compositional abilities -such as combining, interpolating between, or negating instructions -using data alone. To address it, we propose ComposableART, an inference-time technique that extends classifier-free guidance to compositional guidance. It enables the seamless and flexible composition of instructions, leading to highly customizable audio outputs outside the training distribution. Our evaluations across a diverse set of tasks demonstrate that Fugatto performs competitively with specialized models, while ComposableART enhances its sonic palette and control over synthesis. Most notably, we highlight our framework's ability to synthesize emergent sounds -sonic phenomena that transcend conventional audio generation -unlocking new creative possibilities. Demo Website.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- UALM: Unified Audio Language Model for Understanding, Generation and ReasoningJinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh 等ICLR 2026 · 被引用 17 次
- SAO-Instruct: Free-form Audio Editing using Natural Language InstructionsMichael Ungersböck, Florian Grötschla, Luca A. Lanzendörfer, June Young Yi 等NeurIPS 2025 · 被引用 9 次
- AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion ForcingWilliam Chen, Prem Seetharaman, Rithesh Kumar, Oriol Nieto 等ICML 2026 · 被引用 7 次
- AudioStory: Generating Long-Form Narrative Audio with Large Language ModelsYuxin Guo, Teng Wang, Yuying Ge, Shijie Ma 等CVPR 2026 · 被引用 5 次
- SpeechOp: Inference-Time Task Composition for Generative Speech ProcessingJustin Lovelace, Rithesh Kumar, Jiaqi Su, Ke Chen 等ICLR 2026
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Guiding a Diffusion Model with a Bad Version of ItselfTero Karras, Miika Aittala, Tuomas Kynkäänniemi, Jaakko Lehtinen 等NeurIPS 2024 · 被引用 338 次
- Listen, Think, and UnderstandYuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky 等ICLR 2024 · 被引用 247 次
- Compositional Visual Generation with Energy Based ModelsYilun Du, Shuang Li, Igor MordatchNeurIPS 2020 · 被引用 225 次
- Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue AbilitiesZhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping 等ICML 2024 · 被引用 207 次
相关 Paper
- Text-to-Audio Generation using Instruction Guided Latent Diffusion ModelDeepanway Ghosal, Navonil Majumder, Ambuj Mehrish, Soujanya PoriaACM MM 2023 · 被引用 90 次
- BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMsYue Wang, Ruotian Ma, Xingyu Chen, Zhengliang Shi 等ACL 2026
- UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text InstructionsChunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang 等ACL 2026 · 被引用 2 次
- Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic DataSreyan Ghosh, Sonal Kumar, Zhifeng Kong, Rafael Valle 等ICLR 2025
- FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio GenerationYuxuan Jiang, Zehua Chen, Zeqian Ju, Chang Li 等ACM MM 2025 · 被引用 5 次
