Open-Vocabulary Audio-Visual Semantic Segmentation
Ruohao Guo, Liao Qu, Dantong Niu, Yanyu Qi, Wenzhen Yue, Ji Shi, Bowei Xing, Xianghua Ying
Abstract
Audio-visual semantic segmentation (AVSS) aims to segment and classify sounding objects in videos with acoustic cues. However, most approaches operate on the close-set assumption and only identify pre-defined categories from training data, lacking the generalization ability to detect novel categories in practical applications. In this paper, we introduce a new task: open-vocabulary audio-visual semantic segmentation, extending AVSS task to open-world scenarios beyond the annotated label space. This is a more challenging task that requires recognizing all categories, even those that have never been seen nor heard during training. Moreover, we propose the first open-vocabulary AVSS framework, OV-AVSS, which mainly consists of two parts: 1) a universal sound source localization module to perform audio-visual fusion and locate all potential sounding objects and 2) an open-vocabulary classification module to predict categories with the help of the prior knowledge from large-scale pre-trained vision-language models. To properly evaluate the open-vocabulary AVSS, we split zero-shot training and testing subsets based on the AVSBench-semantic benchmark, namely AVSBench-OV. Extensive experiments demonstrate the strong segmentation and zero-shot generalization ability of our model on all categories. On the AVSBench-OV dataset, OV-AVSS achieves 55.43% mIoU on base categories and 29.14% mIoU on novel categories, exceeding the state-of-the-art zero-shot method by 41.88%/20.61% and open-vocabulary method by 10.2%/11.6%. The code is available at https://github.com/ruohaoguo/ovavss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba43c2df-b34a-42ac-9b22-d2da2787bfc1Cited by top-tier papers8
- Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?Jia Li, Wenjie Zhao, Ziru Huang, Yunhui Guo et al.AAAI 2026 · 5 citations
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual SegmentationKaining Ying, Henghui Ding, Guangquan Jie, Yu-Gang JiangICCV 2025 · 3 citations
- Implicit Counterfactual Learning for Audio-Visual SegmentationMingfeng Zha, Tianyu Li, Guoyin Wang, Peng Wang et al.ICCV 2025 · 3 citations
- SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video GenerationKien T. Pham, Yingqing He, Yazhou Xing, Qifeng Chen et al.ACM MM 2025 · 1 citation
- Towards Open-Vocabulary Audio-Visual Event LocalizationJinxing Zhou, Dan Guo, Ruohao Guo, Yuxin Mao et al.CVPR 2025
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- Stepping Out of Similar Semantic Space for Open-Vocabulary SegmentationYong Liu, Song-Li Wu, Sule Bai, Jiahao Wang et al.ICCV 2025 · 6 citations
- Audio-Visual Instance SegmentationRuohao Guo, Xianghua Ying, Yaru Chen, Dantong Niu et al.CVPR 2025
- Unidentified Video Objects: A Benchmark for Dense, Open-World SegmentationWeiyao Wang, Matt Feiszli, Heng Wang, Du TranICCV 2021 · 151 citations
- Audio-Visual Segmentation by Exploring Cross-Modal Mutual SemanticsChen Liu, Peike Patrick Li, Xingqun Qi, Hu Zhang et al.ACM MM 2023 · 33 citations
- Towards Open-Vocabulary Video Instance SegmentationHaochen Wang, Xiaolong Jiang, Xu Tang, Yao Hu et al.ICCV 2023 · 56 citations
