Audio-Visual Segmentation via Unlabeled Frame Exploitation
Jinxiang Liu, Yikun Liu, Fei Zhang, Chen Ju, Ya Zhang, Yanfeng Wang
摘要
Labeled Frame Motion Cues … … Unlabeled Frames Labeled Frame Unlabeled Frames … … Distant Frames GT Supervision No exploitation GT Supervision Semantic Cues (a) Previous methods (w/ GTM) (b) Our proposed method (Ours) (c) Performance comparison Neighboring Frames Figure 1. Comparison between previous methods and ours on how to harness the unlabeled frames. (a) Previous methods perform global temporal modeling (GTM) to process all frames from a sequence including labeled and unlabeled ones, without the exploitation of the unlabeled frames. (b) Our method employs two types of unlabeled frames: (i) the neighboring frames (NFs) provide motion cues for accurately segmenting the sounding object; (ii) the distant frames (DFs) contain semantic cues for enhancing data diversity. (c) Based on TPAVI method, compared to the model trained only using labeled frames (w/o GTM), previous methods using global temporal modeling (w/ GTM) only show marginal performance gain; while our method achieves significant improvement with the unlabeled frames.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced PerformanceHaicheng Wang, Zhemeng Yu, Gabriele Spadaro, Chen Ju 等ICCV 2025 · 被引用 3 次
- Implicit Counterfactual Learning for Audio-Visual SegmentationMingfeng Zha, Tianyu Li, Guoyin Wang, Peng Wang 等ICCV 2025 · 被引用 3 次
- How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?Yujian Lee, Peng Gao, Yongqi Xu, Wentao FanICCV 2025 · 被引用 2 次
- ConText: Driving In-context Learning for Text Removal and SegmentationFei Zhang, Pei Zhang, Baosong Yang, Fei Huang 等ICML 2025
- Video-Guided Foley Sound Generation with Multimodal ControlsZiyang Chen, Prem Seetharaman, Bryan C. Russell, Oriol Nieto 等CVPR 2025
它引用的顶会 Paper26
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- End-to-End Semi-Supervised Object Detection with Soft TeacherMengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang 等ICCV 2021 · 被引用 622 次
- Unbiased Teacher for Semi-Supervised Object DetectionYen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo 等ICLR 2021 · 被引用 603 次
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 被引用 395 次
- PseudoSeg: Designing Pseudo Labels for Semantic SegmentationYuliang Zou, Zizhao Zhang, Han Zhang, Chun-Liang Li 等ICLR 2021 · 被引用 364 次
相关 Paper
- Simultaneously Short- and Long-Term Temporal Modeling for Semi-Supervised Video Semantic SegmentationJiangwei Lao, Weixiang Hong, Xin Guo, Yingying Zhang 等CVPR 2023
- Reducing the Label Bias for Timestamp Supervised Temporal Action SegmentationKaiyuan Liu, Yunheng Li, Shenglan Liu, Chenwei Tan 等CVPR 2023
- Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source LocalizationYuxin Guo, Shijie Ma, Hu Su, Zhiqing Wang 等NeurIPS 2023 · 被引用 19 次
- Iteratively Selecting an Easy Reference Frame Makes Unsupervised Video Object Segmentation EasierYoungjo Lee, Hongje Seong, Euntai KimAAAI 2022 · 被引用 43 次
- Learning Action Completeness from Points for Weakly-supervised Temporal Action LocalizationPilhyeon Lee, Hyeran ByunICCV 2021 · 被引用 81 次
