OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio Separation
Tanvir Mahmud, Diana Marculescu
摘要
Audio separation in real-world scenarios, where mixtures contain a variable number of sources, presents significant challenges due to limitations of existing models, such as over-separation, under-separation, and dependence on predefined training sources. We propose OpenSep, a novel framework that leverages large language models (LLMs) for automated audio separation, eliminating the need for manual intervention and overcoming source limitations. OpenSep uses textual inversion to generate captions from audio mixtures with off-the-shelf audio captioning models, effectively parsing the sound sources present. It then employs few-shot LLM prompting to extract detailed audio properties of each parsed source, facilitating separation in unseen mixtures. Additionally, we introduce a multi-level extension of the mix-and-separate training framework to enhance modality alignment by separating single source sounds and mixtures simultaneously. Extensive experiments demonstrate OpenSep’s superiority in precisely separating new, unseen, and variable sources in challenging mixtures, outperforming SOTA baseline methods. Code is released at https://github.com/tanvir-utexas/OpenSep.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen SoundsEfthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey 等ICLR 2021 · 被引用 83 次
- Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataKe Chen, Xingjian Du, Bilei Zhu, Zejun Ma 等AAAI 2022 · 被引用 58 次
- Visual Scene Graphs for Audio Source SeparationMoitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja, Anoop CherianICCV 2021 · 被引用 45 次
- Weakly-supervised Audio Separation via Bi-modal Semantic SimilarityTanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida, Diana MarculescuICLR 2024 · 被引用 4 次
相关 Paper
- ZeroSep: Separate Anything in Audio with Zero TrainingChao Huang, Yuesheng Ma, Junxuan Huang, Susan Liang 等NeurIPS 2025 · 被引用 8 次
- Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio EncodersWeiqiao Shan, Yuang Li, Yuhao Zhang, Yingfeng Luo 等EMNLP 2025 · 被引用 1 次
- MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic AlignmentHao Zhou, Xiaobao Guo, Yuzhe Zhu, Adams Wai-Kin KongAAAI 2026 · 被引用 2 次
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio CaptioningSangyeon Cho, Mingi Kim, Jinkwon Hwang, Jaehoon Go 等EMNLP 2025
- Language-Guided Audio-Visual Source Separation via Trimodal ConsistencyReuben Tan, Arijit Ray, Andrea Burns, Bryan A. Plummer 等CVPR 2023
