Implicit Counterfactual Learning for Audio-Visual Segmentation
Mingfeng Zha, Tianyu Li, Guoyin Wang, Peng Wang, Yangyang Wu, Yang Yang, Heng Tao Shen
摘要
Audio-visual segmentation (AVS) aims to segment objects in videos based on audio cues. Existing AVS methods are primarily designed to enhance interaction efficiency but pay limited attention to modality representation discrepancies and imbalances. To overcome this, we propose the implicit counterfactual framework (ICF) to achieve unbiased cross-modal understanding. Due to the lack of semantics, heterogeneous representations may lead to erroneous matches, especially in complex scenes with ambiguous visual content or interference from multiple audio sources. We introduce the multi-granularity implicit text (MIT) involving video-, segment- and frame-level as the bridge to establish the modality-shared space, reducing modality gaps and providing prior guidance. Visual content carries more information and typically dominates, thereby marginalizing audio features in the decision-making. To mitigate knowledge preference, we propose the semantic counterfactual (SC) to learn orthogonal representations in the latent space, generating diverse counterfactual samples, thus avoiding biases introduced by complex functional designs and explicit modifications of text structures or attributes. We further formulate the collaborative distribution-aware contrastive learning (CDCL), incorporating factual-counterfactual and inter-modality contrasts to align representations, promoting cohesion and decoupling. Extensive experiments on three public datasets validate that the proposed method achieves state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AVTrack: Audio-Visual Tracking in Human-centric Complex ScenesYaoting Wang, Yun Zhou, Zipei Zhang, Henghui DingICML 2026 · 被引用 1 次
- Bootstrap Your Own AV-Proxies: Adaptive Contrastive and Prototype Learning for Audio-Visual SegmentationJunbo Zhang, Hang Su, Zhaofan Li, Hang Dong 等CVPR 2026
它引用的顶会 Paper45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent AlignmentChen Liu, Peike Li, Liying Yang, Dadong Wang 等CVPR 2025
- TAVT: Towards Transferable Audio-Visual Text GenerationWang Lin, Tao Jin, Wenwen Pan, Linjun Li 等ACL 2023 · 被引用 9 次
- Counterfactual Contrastive Learning for Weakly-Supervised Vision-Language GroundingZhu Zhang, Zhou Zhao, Zhijie Lin, Jieming Zhu 等NeurIPS 2020 · 被引用 74 次
- Distilling Audio-Visual Knowledge by Compositional Contrastive LearningYanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan 等CVPR 2021
- SCLAV: Supervised Cross-modal Contrastive Learning for Audio-Visual CodingChao Sun, Min Chen, Jialiang Cheng, Han Liang 等ACM MM 2023 · 被引用 3 次
