Weakly-Supervised Audio-Visual Segmentation
Shentong Mo, Bhiksha Raj
Abstract
Audio-visual segmentation is a challenging task that aims to predict pixel-level masks for sound sources in a video. Previous work applied a comprehensive manually designed architecture with countless pixel-wise accurate masks as supervision. However, these pixel-level masks are expensive and not available in all cases. In this work, we aim to simplify the supervision as the instance-level annotation, i.e., weakly-supervised audio-visual segmentation. We present a novel Weakly-Supervised Audio-Visual Segmentation framework, namely WS-AVS, that can learn multi-scale audio-visual alignment with multi-scale multiple-instance contrastive learning for audio-visual segmentation. Extensive experiments on AVS-Bench demonstrate the effectiveness of our WS-AVS in the weakly-supervised audio-visual segmentation of single-source and multi-source scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object SegmentationShaofei Huang, Rui Ling, Hongyu Li, Tianrui Hui et al.AAAI 2025 · 24 citations
- Aligning Audio-Visual Joint Representations with an Agentic WorkflowShentong Mo, Yibing SongNeurIPS 2024 · 7 citations
- Unveiling and Mitigating Bias in Audio Visual SegmentationPeiwen Sun, Honggang Zhang, Di HuACM MM 2024 · 7 citations
- Implicit Counterfactual Learning for Audio-Visual SegmentationMingfeng Zha, Tianyu Li, Guoyin Wang, Peng Wang et al.ICCV 2025 · 3 citations
- CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-TrainingYuxin Guo, Siyang Sun, Shuailei Ma, Kecheng Zheng et al.CVPR 2024
Builds on24
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
- Learning Representations from Audio-Visual Spatial AlignmentPedro Morgado, Yi Li, Nuno VasconcelosNeurIPS 2020 · 149 citations
- Self-Supervised Difference Detection for Weakly-Supervised Semantic SegmentationWataru Shimoda, Keiji YanaiICCV 2019 · 148 citations
- C2 AM: Contrastive learning of Class-agnostic Activation Map for Weakly Supervised Object Localization and Semantic SegmentationJinheng Xie, Jianfeng Xiang, Junliang Chen, Xianxu Hou et al.CVPR 2022 · 139 citations
Related papers
- Unraveling Instance Associations: A Closer Look for Audio-Visual SegmentationYuanhong Chen, Yuyuan Liu, Hu Wang, Fengbei Liu et al.CVPR 2024
- Unsupervised Sounding Pixel LearningYining Zhang, Yanli Ji, Yang YangEMNLP 2023 · 2 citations
- Audio-Visual Segmentation by Exploring Cross-Modal Mutual SemanticsChen Liu, Peike Patrick Li, Xingqun Qi, Hu Zhang et al.ACM MM 2023 · 33 citations
- Weakly-Supervised Audio-Visual Video Parsing with Prototype-Based Pseudo-LabelingKranthi Kumar Rachavarapu, Kalyan Ramakrishnan, A. N. RajagopalanCVPR 2024 · 3 citations
- Unsupervised Audio-Visual Segmentation with Modality AlignmentSwapnil Bhosale, Haosen Yang, Diptesh Kanojia, Jiankang Deng et al.AAAI 2025 · 11 citations
