Implicit Counterfactual Learning for Audio-Visual Segmentation
Mingfeng Zha, Tianyu Li, Guoyin Wang, Peng Wang, Yangyang Wu, Yang Yang, Heng Tao Shen
Abstract
Audio-visual segmentation (AVS) aims to segment objects in videos based on audio cues. Existing AVS methods are primarily designed to enhance interaction efficiency but pay limited attention to modality representation discrepancies and imbalances. To overcome this, we propose the implicit counterfactual framework (ICF) to achieve unbiased cross-modal understanding. Due to the lack of semantics, heterogeneous representations may lead to erroneous matches, especially in complex scenes with ambiguous visual content or interference from multiple audio sources. We introduce the multi-granularity implicit text (MIT) involving video-, segment- and frame-level as the bridge to establish the modality-shared space, reducing modality gaps and providing prior guidance. Visual content carries more information and typically dominates, thereby marginalizing audio features in the decision-making. To mitigate knowledge preference, we propose the semantic counterfactual (SC) to learn orthogonal representations in the latent space, generating diverse counterfactual samples, thus avoiding biases introduced by complex functional designs and explicit modifications of text structures or attributes. We further formulate the collaborative distribution-aware contrastive learning (CDCL), incorporating factual-counterfactual and inter-modality contrasts to align representations, promoting cohesion and decoupling. Extensive experiments on three public datasets validate that the proposed method achieves state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98cb15e0-14e0-484a-87e7-b5c830c8eb3bCited by top-tier papers2
- AVTrack: Audio-Visual Tracking in Human-centric Complex ScenesYaoting Wang, Yun Zhou, Zipei Zhang, Henghui DingICML 2026 · 1 citation
- Bootstrap Your Own AV-Proxies: Adaptive Contrastive and Prototype Learning for Audio-Visual SegmentationJunbo Zhang, Hang Su, Zhaofan Li, Hang Dong et al.CVPR 2026
Builds on45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent AlignmentChen Liu, Peike Li, Liying Yang, Dadong Wang et al.CVPR 2025
- TAVT: Towards Transferable Audio-Visual Text GenerationWang Lin, Tao Jin, Wenwen Pan, Linjun Li et al.ACL 2023 · 9 citations
- Counterfactual Contrastive Learning for Weakly-Supervised Vision-Language GroundingZhu Zhang, Zhou Zhao, Zhijie Lin, Jieming Zhu et al.NeurIPS 2020 · 74 citations
- Distilling Audio-Visual Knowledge by Compositional Contrastive LearningYanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan et al.CVPR 2021
- SCLAV: Supervised Cross-modal Contrastive Learning for Audio-Visual CodingChao Sun, Min Chen, Jialiang Cheng, Han Liang et al.ACM MM 2023 · 3 citations
