Audio-Visual Segmentation via Unlabeled Frame Exploitation
Jinxiang Liu, Yikun Liu, Fei Zhang, Chen Ju, Ya Zhang, Yanfeng Wang
Abstract
Labeled Frame Motion Cues … … Unlabeled Frames Labeled Frame Unlabeled Frames … … Distant Frames GT Supervision No exploitation GT Supervision Semantic Cues (a) Previous methods (w/ GTM) (b) Our proposed method (Ours) (c) Performance comparison Neighboring Frames Figure 1. Comparison between previous methods and ours on how to harness the unlabeled frames. (a) Previous methods perform global temporal modeling (GTM) to process all frames from a sequence including labeled and unlabeled ones, without the exploitation of the unlabeled frames. (b) Our method employs two types of unlabeled frames: (i) the neighboring frames (NFs) provide motion cues for accurately segmenting the sounding object; (ii) the distant frames (DFs) contain semantic cues for enhancing data diversity. (c) Based on TPAVI method, compared to the model trained only using labeled frames (w/o GTM), previous methods using global temporal modeling (w/ GTM) only show marginal performance gain; while our method achieves significant improvement with the unlabeled frames.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced PerformanceHaicheng Wang, Zhemeng Yu, Gabriele Spadaro, Chen Ju et al.ICCV 2025 · 3 citations
- Implicit Counterfactual Learning for Audio-Visual SegmentationMingfeng Zha, Tianyu Li, Guoyin Wang, Peng Wang et al.ICCV 2025 · 3 citations
- How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?Yujian Lee, Peng Gao, Yongqi Xu, Wentao FanICCV 2025 · 2 citations
- ConText: Driving In-context Learning for Text Removal and SegmentationFei Zhang, Pei Zhang, Baosong Yang, Fei Huang et al.ICML 2025
- Video-Guided Foley Sound Generation with Multimodal ControlsZiyang Chen, Prem Seetharaman, Bryan C. Russell, Oriol Nieto et al.CVPR 2025
Builds on26
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- End-to-End Semi-Supervised Object Detection with Soft TeacherMengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang et al.ICCV 2021 · 622 citations
- Unbiased Teacher for Semi-Supervised Object DetectionYen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo et al.ICLR 2021 · 603 citations
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 395 citations
- PseudoSeg: Designing Pseudo Labels for Semantic SegmentationYuliang Zou, Zizhao Zhang, Han Zhang, Chun-Liang Li et al.ICLR 2021 · 364 citations
Related papers
- Simultaneously Short- and Long-Term Temporal Modeling for Semi-Supervised Video Semantic SegmentationJiangwei Lao, Weixiang Hong, Xin Guo, Yingying Zhang et al.CVPR 2023
- Reducing the Label Bias for Timestamp Supervised Temporal Action SegmentationKaiyuan Liu, Yunheng Li, Shenglan Liu, Chenwei Tan et al.CVPR 2023
- Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source LocalizationYuxin Guo, Shijie Ma, Hu Su, Zhiqing Wang et al.NeurIPS 2023 · 19 citations
- Iteratively Selecting an Easy Reference Frame Makes Unsupervised Video Object Segmentation EasierYoungjo Lee, Hongje Seong, Euntai KimAAAI 2022 · 43 citations
- Learning Action Completeness from Points for Weakly-supervised Temporal Action LocalizationPilhyeon Lee, Hyeran ByunICCV 2021 · 81 citations
