SAM Audio: Segment Anything in Audio
Bowen Shi, Andros Tjandra, John Hoffman, Helin Wang, YI-CHIAO WU, Luya Gao, Julius Richter, Matthew Le, Apoorv Vyas, Sanyuan Chen, Christoph Feichtenhofer, Piotr Dollár
Abstract
General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separation models are either domain-specific, designed for fixed categories such as speech or music, or limited in controllability, supporting only a single prompting modality such as text. In this work, we present SAM AUDIO, a foundation model for general audio separation that unifies text, visual, and temporal span prompting within a single framework. Built on a diffusion transformer architecture, SAM AUDIO is trained with flow matching on large-scale audio data spanning speech, music, and general sounds, and can flexibly separate target sources described by language, visual masks, or temporal spans. The model achieves state-of-the-art performance across a diverse suite of benchmarks, including general sound, speech, music, and musical instrument separation in both in-the-wild and professionally produced audios, substantially outperforming prior general-purpose and specialized systems. Furthermore, we introduce a new real-world separation benchmark with human-labeled multimodal prompts and a reference-free evaluation model that correlates strongly with human judgment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar et al.NeurIPS 2023 · 910 citations
- Voicebox: Text-Guided Multilingual Universal Speech Generation at ScaleMatthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer et al.NeurIPS 2023 · 613 citations
Related papers
- ZeroSep: Separate Anything in Audio with Zero TrainingChao Huang, Yuesheng Ma, Junxuan Huang, Susan Liang et al.NeurIPS 2025 · 8 citations
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren et al.ICML 2023 · 469 citations
- Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?Jia Li, Wenjie Zhao, Ziru Huang, Yunhui Guo et al.AAAI 2026 · 5 citations
- TSAM: Temporal SAM Augmented with Multimodal Prompts for Referring Audio-Visual SegmentationAbduljalil Radman, Jorma LaaksonenCVPR 2025
- Language-Guided Audio-Visual Source Separation via Trimodal ConsistencyReuben Tan, Arijit Ray, Andrea Burns, Bryan A. Plummer et al.CVPR 2023
