Improving Audio-Visual Segmentation with Bidirectional Generation
Dawei Hao, Yuxin Mao, Bowen He, Xiaodong Han, Yuchao Dai, Yiran Zhong
Abstract
The aim of audio-visual segmentation (AVS) is to precisely differentiate audible objects within videos down to the pixel level. Traditional approaches often tackle this challenge by combining information from various modalities, where the contribution of each modality is implicitly or explicitly modeled. Nevertheless, the interconnections between different modalities tend to be overlooked in audio-visual modeling. In this paper, inspired by the human ability to mentally simulate the sound of an object and its visual appearance, we introduce a bidirectional generation framework. This framework establishes robust correlations between an object's visual characteristics and its associated sound, thereby enhancing the performance of AVS. To achieve this, we employ a visual-to-audio projection component that reconstructs audio features from object segmentation masks and minimizes reconstruction errors. Moreover, recognizing that many sounds are linked to object movements, we introduce an implicit volumetric motion estimation module to handle temporal dynamics that may be challenging to capture using conventional optical flow methods. To showcase the effectiveness of our approach, we conduct comprehensive experiments and analyses on the widely recognized AVSBench benchmark. As a result, we establish a new state-of-the-art performance level in the AVS benchmark, particularly excelling in the challenging MS3 subset which involves segmenting multiple sound sources. Code is released in: https://github.com/ OpenNLPLab/AVS-bidirectional.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 499a7944-f4f9-44ee-9cb7-13b63fe7d072Cited by top-tier papers17
- Mixtures of Experts for Audio-Visual LearningYing Cheng, Yang Li, Junjie He, Rui FengNeurIPS 2024 · 22 citations
- Exploring Transformer ExtrapolationZhen Qin, Yiran Zhong, Hui DengAAAI 2024 · 12 citations
- Unsupervised Audio-Visual Segmentation with Modality AlignmentSwapnil Bhosale, Haosen Yang, Diptesh Kanojia, Jiankang Deng et al.AAAI 2025 · 11 citations
- Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual SegmentationQi Yang, Xing Nie, Tong Li, Pengfei Gao et al.CVPR 2024 · 9 citations
- Unveiling and Mitigating Bias in Audio Visual SegmentationPeiwen Sun, Honggang Zhang, Di HuACM MM 2024 · 7 citations
Builds on12
- Self-supervised Video Object Segmentation by Motion GroupingCharig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman et al.ICCV 2021 · 188 citations
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan et al.NeurIPS 2020 · 156 citations
- Learning Generative Vision Transformer with Energy-Based Latent Space for Saliency PredictionJing Zhang, Jianwen Xie, Nick Barnes, Ping LiNeurIPS 2021 · 117 citations
- Cross-Modal Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationHanyu Xuan, Zhenyu Zhang, Shuo Chen, Jian Yang et al.AAAI 2020 · 110 citations
- Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningYing Cheng, Ruize Wang, Zhihao Pan, Rui Feng et al.ACM MM 2020 · 93 citations
Related papers
- Multimodal Variational Auto-encoder based Audio-Visual SegmentationYuxin Mao, Jing Zhang, Mochu Xiang, Yiran Zhong et al.ICCV 2023 · 57 citations
- Audio-Visual Segmentation by Exploring Cross-Modal Mutual SemanticsChen Liu, Peike Patrick Li, Xingqun Qi, Hu Zhang et al.ACM MM 2023 · 33 citations
- Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?Jia Li, Wenjie Zhao, Ziru Huang, Yunhui Guo et al.AAAI 2026 · 5 citations
- AVSegFormer: Audio-Visual Segmentation with TransformerShengyi Gao, Zhe Chen, Guo Chen, Wenhai Wang et al.AAAI 2024 · 96 citations
- Audio-Visual Instance SegmentationRuohao Guo, Xianghua Ying, Yaru Chen, Dantong Niu et al.CVPR 2025
