Weakly Supervised Instance Segmentation for Videos With Temporal Mask Consistency
Qing Liu, Vignesh Ramanathan, Dhruv Mahajan, Alan L. Yuille, Zhenheng Yang
Abstract
Weakly supervised instance segmentation reduces the cost of annotations required to train models. However, existing approaches which rely only on image-level class labels predominantly suffer from errors due to (a) partial segmentation of objects and (b) missing object predictions. We show that these issues can be better addressed by training with weakly labeled videos instead of images. In videos, motion and temporal consistency of predictions across frames provide complementary signals which can help segmentation. We are the first to explore the use of these video signals to tackle weakly supervised instance segmentation. We propose two ways to leverage this information in our model. First, we adapt inter-pixel relation network (IRN) [1] to effectively incorporate motion information during training. Second, we introduce a new MaskConsist module, which addresses the problem of missing object instances by transferring stable predictions between neighboring frames during training. We demonstrate that both approaches together improve the instance segmentation metric AP 50 on video frames of two datasets: Youtube-VIS and Cityscapes by 5% and 3% respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b9a8fb0-cdd3-40f8-b874-c4f697477305Cited by top-tier papers4
- MinVIS: A Minimal Video Instance Segmentation Framework without Video-based TrainingDe-An Huang, Zhiding Yu, Anima AnandkumarNeurIPS 2022 · 135 citations
- 2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level SupervisionCheng-Kun Yang, Min-Hung Chen, Yung-Yu Chuang, Yen-Yu LinICCV 2023 · 30 citations
- Mask-Free Video Instance SegmentationLei Ke, Martin Danelljan, Henghui Ding, Yu-Wing Tai et al.CVPR 2023
- Hierarchical Visual Prompt Learning for Continual Video Instance SegmentationJiahua Dong, Hui Yin, Wenqi Liang, Hanbin Zhao et al.ICCV 2025
Builds on7
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Self-Supervised Difference Detection for Weakly-Supervised Semantic SegmentationWataru Shimoda, Keiji YanaiICCV 2019 · 148 citations
- Label-PEnet: Sequential Label Propagation and Enhancement Networks for Weakly Supervised Instance SegmentationWeifeng Ge, Weilin Huang, Sheng Guo, Matthew R. ScottICCV 2019 · 54 citations
- Frame-to-Frame Aggregation of Active Regions in Web Videos for Weakly Supervised Semantic SegmentationJungbeom Lee, Eunji Kim, Sungmin Lee, Jangho Lee et al.ICCV 2019 · 45 citations
- Classifying, Segmenting, and Tracking Object Instances in Video with Mask PropagationGedas Bertasius, Lorenzo TorresaniCVPR 2020
Related papers
- Learning to Track Instances without Video AnnotationsYang Fu, Sifei Liu, Umar Iqbal, Shalini De Mello et al.CVPR 2021
- Minimizing Labeled, Maximizing Unlabeled: An Image-Driven Approach for Video Instance SegmentationFangyun Wei, Jinjing Zhao, Kun Yan, Chang XuCVPR 2025
- Class-incremental Continual Learning for Instance Segmentation with Image-level Weak SupervisionYu-Hsing Hsieh, Guan-Sheng Chen, Shun-Xian Cai, Ting-Yun Wei et al.ICCV 2023 · 16 citations
- InstMove: Instance Motion for Object-centric Video SegmentationQihao Liu, Junfeng Wu, Yi Jiang, Xiang Bai et al.CVPR 2023
- Self-Supervised Multi-Object Tracking with Cross-input ConsistencyFavyen Bastani, Songtao He, Samuel MaddenNeurIPS 2021 · 39 citations
