Two-shot Video Object Segmentation
Kun Yan, Xiao Li, Fangyun Wei, Jinglu Wang, Chenbin Zhang, Ping Wang, Yan Lu
Abstract
Previous works on video object segmentation (VOS) are trained on densely annotated videos. Nevertheless, acquiring annotations in pixel level is expensive and timeconsuming. In this work, we demonstrate the feasibility of training a satisfactory VOS model on sparsely annotated videos-we merely require two labeled frames per training video while the performance is sustained. We term this novel training paradigm as two-shot video object segmentation, or two-shot VOS for short. The underlying idea is to generate pseudo labels for unlabeled frames during training and to optimize the model on the combination of labeled and pseudo-labeled data. Our approach is extremely simple and can be applied to a majority of existing frameworks. We first pre-train a VOS model on sparsely annotated videos in a semi-supervised manner, with the first frame always being a labeled one. Then, we adopt the pretrained VOS model to generate pseudo labels for all unlabeled frames, which are subsequently stored in a pseudolabel bank. Finally, we retrain a VOS model on both labeled and pseudo-labeled data without any restrictions on the first frame. For the first time, we present a general way to train VOS models on two-shot VOS datasets. By using 7.3% and 2.9% labeled data of YouTube-VOS and DAVIS benchmarks, our approach achieves comparable results in contrast to the counterparts trained on fully labeled set. Code and models are available at https://github.com/yk- pku/Two-shot-Video-Object-Segmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and TextXiang Li, Jinglu Wang, Xiaohao Xu, Muqiao Yang et al.EMNLP 2023 · 6 citations
- EVOLVE: Event-Guided Deformable Feature Transfer and Dual-Memory Refinement for Low-Light Video Object SegmentationJong-Hyeon Baek, Jiwon Oh, Yeong Jun KohICCV 2025 · 2 citations
- Putting the Object Back into Video Object SegmentationHo Kei Cheng, Seoung Wug Oh, Brian L. Price, Joon-Young Lee et al.CVPR 2024
- Minimizing Labeled, Maximizing Unlabeled: An Image-Driven Approach for Video Instance SegmentationFangyun Wei, Jinjing Zhao, Kun Yan, Chang XuCVPR 2025
Builds on24
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- End-to-End Semi-Supervised Object Detection with Soft TeacherMengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang et al.ICCV 2021 · 622 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
Related papers
- Learning Fast and Robust Target Models for Video Object SegmentationAndreas Robinson, Felix Järemo Lawin, Martin Danelljan, Fahad Shahbaz Khan et al.CVPR 2020
- From ViT Features to Training-free Video Object Segmentation via Streaming-data Mixture ModelsRoy Uziel, Or Dinari, Oren FreifeldNeurIPS 2023 · 6 citations
- Unified Mask Embedding and Correspondence Learning for Self-Supervised Video SegmentationLiulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li et al.CVPR 2023
- Make One-Shot Video Object Segmentation Efficient AgainTim Meinhardt, Laura Leal-TaixéNeurIPS 2020 · 44 citations
- Learning Video Object Segmentation From Unlabeled VideosXiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai et al.CVPR 2020
