OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
Xize Cheng, Siqi Zheng, Zehan Wang, Minghui Fang, Ziang Zhang, Rongjie Huang, Shengpeng Ji, Jialong Zuo, Tao Jin, Zhou Zhao
Abstract
The scaling up has brought tremendous success in the fields of vision and language in recent years. When it comes to audio, however, researchers encounter a major challenge in scaling up the training data, as most natural audio contains diverse interfering signals. To address this limitation, we introduce Omni-modal Sound Separation (OmniSep), a novel framework capable of isolating clean soundtracks based on omni-modal queries, encompassing both single-modal and multi-modal composed queries. Specifically, we introduce the Query-Mixup strategy, which blends query features from different modalities during training. This enables OmniSep to optimize multiple modalities concurrently, effectively bringing all modalities under a unified framework for sound separation. We further enhance this flexibility by allowing queries to influence sound separation positively or negatively, facilitating the retention or removal of specific sounds as desired. Finally, OmniSep employs a retrieval-augmented approach known as Query-Aug, which enables open-vocabulary sound separation. Experimental evaluations on MUSIC, VGGSOUND-CLEAN+, and MUSIC-CLEAN+ datasets demonstrate effectiveness of OmniSep, achieving state-of-the-art performance in text-, image-, and audio-queried sound separation tasks. For samples and further information, please visit the demo page at https://omnisep.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64e80e16-a14d-4e00-9d4a-f6a4dc514c57Cited by top-tier papers6
- Audio-sync Video Instance Editing with Granularity-Aware Mask RefinerHaojie Zheng, Shuchen Weng, Jingqi Liu, Siqi Yang et al.CVPR 2026 · 8 citations
- A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound SeparationKai Li, Jintao Cheng, Chang Zeng, Zijun Yan et al.ICML 2026 · 2 citations
- MARS-Sep: Multimodal-Aligned Reinforced Sound SeparationZihan Zhang, Xize Cheng, Zhennan Jiang, Dongjie Fu et al.ICLR 2026 · 2 citations
- Cinematic Audio Source Separation Using Visual CuesKang Zhang, Suyeon Lee, Arda Senocak, Joon Son ChungCVPR 2026 · 1 citation
- USE: A Unified Model for Universal Sound Separation and ExtractionHongyu Wang, Chenda Li, Xin Zhou, Shuai Wang et al.AAAI 2026
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan et al.ICLR 2024 · 403 citations
Related papers
- OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio SeparationTanvir Mahmud, Diana MarculescuEMNLP 2024 · 1 citation
- AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow MatchingXize Cheng, Chenyuhao Wen, Slytherin Wang, Yongqi Wang et al.ICLR 2026
- Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual ScenesHyeonggon Ryu, Seongyu Kim, Joon Son Chung, Arda SenocakCVPR 2025
- CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled VideosHao-Wen Dong, Naoya Takahashi, Yuki Mitsufuji, Julian J. McAuley et al.ICLR 2023 · 3 citations
- Weakly-supervised Audio Separation via Bi-modal Semantic SimilarityTanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida, Diana MarculescuICLR 2024 · 4 citations
