Iteratively Selecting an Easy Reference Frame Makes Unsupervised Video Object Segmentation Easier
Youngjo Lee, Hongje Seong, Euntai Kim
Abstract
Unsupervised video object segmentation (UVOS) is a per-pixel binary labeling problem which aims at separating the foreground object from the background in the video without using the ground truth (GT) mask of the foreground object. Most of the previous UVOS models use the first frame or the entire video as a reference frame to specify the mask of the foreground object. Our question is why the first frame should be selected as a reference frame or why the entire video should be used to specify the mask. We believe that we can select a better reference frame to achieve the better UVOS performance than using only the first frame or the entire video as a reference frame. In our paper, we propose Easy Frame Selector (EFS). The EFS enables us to select an "easy" reference frame that makes the subsequent VOS become easy, thereby improving the VOS performance. Furthermore, we propose a new framework named as Iterative Mask Prediction (IMP). In the framework, we repeat applying EFS to the given video and selecting an "easier" reference frame from the video than the previous iteration, increasing the VOS performance incrementally. The IMP consists of EFS, Bi-directional Mask Prediction (BMP), and Temporal Information Updating (TIU). From the proposed framework, we achieve state-of-the-art performance in three UVOS benchmark sets: DAVIS16, FBMS, and SegTrack-V2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Weakly Supervised Video Salient Object Detection via Point SupervisionShuyong Gao, Haozhe Xing, Wei Zhang, Yan Wang et al.ACM MM 2022 · 39 citations
- Unsupervised Video Object Segmentation with Online Adversarial Self-TuningTiankang Su, Huihui Song, Dong Liu, Bo Liu et al.ICCV 2023 · 19 citations
- Timeline and Boundary Guided Diffusion Network for Video Shadow DetectionHaipeng Zhou, Hongqiu Wang, Tian Ye, Zhaohu Xing et al.ACM MM 2024 · 18 citations
- Isomer: Isomerous Transformer for Zero-shot Video Object SegmentationYichen Yuan, Yifan Wang, Lijun Wang, Xiaoqi Zhao et al.ICCV 2023 · 16 citations
- SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationLingyi Hong, Wei Zhang, Shuyong Gao, Hong Lu et al.ACM MM 2023 · 14 citations
Builds on6
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Stacked Cross Refinement Network for Edge-Aware Salient Object DetectionZhe Wu, Li Su, Qingming HuangICCV 2019 · 374 citations
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao et al.AAAI 2020 · 210 citations
- Pyramid Constrained Self-Attention Network for Fast Video Salient Object DetectionYuchao Gu, Lijuan Wang, Ziqin Wang, Yun Liu et al.AAAI 2020 · 184 citations
- Anchor Diffusion for Unsupervised Video Object SegmentationZhao Yang, Qiang Wang, Luca Bertinetto, Song Bai et al.ICCV 2019 · 127 citations
Related papers
- Unified Mask Embedding and Correspondence Learning for Self-Supervised Video SegmentationLiulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li et al.CVPR 2023
- FlowTrack: Integrating Adjacent-Frame Motion Tracking and Adaptive Prediction for Robust Semi-Supervised VOSDuolin Wang, Guanyu Xing, Yanli LiuACM MM 2025 · 1 citation
- Learning Video Object Segmentation From Unlabeled VideosXiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai et al.CVPR 2020
- Integrating Boxes and Masks: A Multi-Object Framework for Unified Visual Tracking and SegmentationYuanyou Xu, Zongxin Yang, Yi YangICCV 2023 · 18 citations
- Two-shot Video Object SegmentationKun Yan, Xiao Li, Fangyun Wei, Jinglu Wang et al.CVPR 2023
