Audio-Visual Segmentation by Exploring Cross-Modal Mutual Semantics
Chen Liu, Peike Patrick Li, Xingqun Qi, Hu Zhang, Lincheng Li, Dadong Wang, Xin Yu
Abstract
The audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we observed that prior arts are prone to segment a certain salient object in a video regardless of the audio information. This is because sounding objects are often the most salient ones in the AVS dataset. Thus, current AVS methods might fail to localize genuine sounding objects due to the dataset bias. In this work, we present an audio-visual instance-aware segmentation approach to overcome the dataset bias. In a nutshell, our method first localizes potential sounding objects in a video by an object segmentation network, and then associates the sounding object candidates with the given audio. We notice that an object could be a sounding object in one video but a silent one in another video. This would bring ambiguity in training our object segmentation network as only sounding objects have corresponding segmentation masks. We thus propose a silent object-aware segmentation objective to alleviate the ambiguity. Moreover, since the category information of audio is unknown, especially for multiple sounding sources, we propose to explore the audio-visual semantic correlation and then associate audio with potential objects. Specifically, we attend predicted audio category scores to potential instance masks and these scores will highlight corresponding sounding instances while suppressing inaudible ones. When we enforce the attended instance masks to resemble the ground-truth mask, we are able to establish audio-visual semantics correlation. Experimental results on the AVS benchmarks demonstrate that our method can effectively segment sounding objects without being biased to salient objects and also achieves state-of-the-art performance in both the single-source and multi-source scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec6930af-4f99-4916-9314-a4b72a0154ceCited by top-tier papers17
- Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-Speech Gesture GenerationXingqun Qi, Jiahao Pan, Peng Li, Ruibin Yuan et al.CVPR 2024 · 9 citations
- Unveiling and Mitigating Bias in Audio Visual SegmentationPeiwen Sun, Honggang Zhang, Di HuACM MM 2024 · 7 citations
- Dense Audio-Visual Event Localization Under Cross-Modal Consistency and Multi-Temporal Granularity CollaborationZiheng Zhou, Jinxing Zhou, Wei Qian, Shengeng Tang et al.AAAI 2025 · 4 citations
- Open-Vocabulary Audio-Visual Semantic SegmentationRuohao Guo, Liao Qu, Dantong Niu, Yanyu Qi et al.ACM MM 2024 · 4 citations
- Benchmarking Audio Visual Segmentation for Long-Untrimmed VideosChen Liu, Peike Patrick Li, Qingtao Yu, Hongwei Sheng et al.CVPR 2024
Builds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 1,898 citations
- BEATs: Audio Pre-Training with Acoustic TokenizersSanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu et al.ICML 2023 · 568 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
Related papers
- Audio-Visual Instance SegmentationRuohao Guo, Xianghua Ying, Yaru Chen, Dantong Niu et al.CVPR 2025
- Weakly-Supervised Audio-Visual SegmentationShentong Mo, Bhiksha RajNeurIPS 2023 · 26 citations
- Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?Jia Li, Wenjie Zhao, Ziru Huang, Yunhui Guo et al.AAAI 2026 · 5 citations
- Audio-Visual Grouping Network for Sound Localization from MixturesShentong Mo, Yapeng TianCVPR 2023
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan et al.NeurIPS 2020 · 156 citations
