CTVIS: Consistent Training for Online Video Instance Segmentation
Kaining Ying, Qing Zhong, Weian Mao, Zhenhua Wang, Hao Chen, Lin Yuanbo Wu, Yifan Liu, Chengxiang Fan, Yunzhi Zhuge, Chunhua Shen
Abstract
The discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/negative embeddings. Recent online VIS methods leverage CIs sourced from one reference frame only, which we argue is insufficient for learning highly discriminative embeddings. Intuitively, a possible strategy to enhance CIs is replicating the inference phase during training. To this end, we propose a simple yet effective training strategy, called Consistent Training for Online VIS (CTVIS), which devotes to aligning the training and inference pipelines in terms of building CIs. Specifically, CTVIS constructs CIs by referring inference the momentum-averaged embedding and the memory bank storage mechanisms, and adding noise to the relevant embeddings. Such an extension allows a reliable comparison between embeddings of current instances and the stable representations of historical instances, thereby conferring an advantage in modeling VIS challenges such as occlusion, re-identification, and deformation. Empirically, CTVIS outstrips the SOTA VIS models by up to +5.0 points on three VIS benchmarks, including YTVIS19 (55.1% AP), YTVIS21 (50.1% AP) and OVIS (35.5% AP). Furthermore, we find that pseudo-videos transformed from images can train robust models surpassing fully-supervised ones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f497fb0-225a-4f21-9e07-9aab974612a9Cited by top-tier papers19
- OnlineTAS: An Online Baseline for Temporal Action SegmentationQing Zhong, Guodong Ding, Angela YaoNeurIPS 2024 · 15 citations
- SyncVIS: Synchronized Video Instance SegmentationRongkun Zheng, Lu Qi, Xi Chen, Yi Wang et al.NeurIPS 2024 · 8 citations
- ViLLa: Video Reasoning Segmentation with Large Language ModelRongkun Zheng, Lu Qi, Xi Chen, Yi Wang et al.ICCV 2025 · 7 citations
- The STVchrono Dataset: Towards Continuous Change Recognition in TimeYanjun Sun, Yue Qiu, Mariia Khan, Fumiya Matsuzawa et al.CVPR 2024 · 6 citations
- OnlineHMR: Video-based Online World-Grounded Human Mesh RecoveryYiwen Zhao, Ce Zheng, Yufu Wang, Hsueh-Han Daniel Yang et al.CVPR 2026 · 5 citations
Builds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Video Instance Segmentation using Inter-Frame Communication TransformersSukjun Hwang, Miran Heo, Seoung Wug Oh, Seon Joo KimNeurIPS 2021 · 174 citations
Related papers
- Crossover Learning for Fast Online Video Instance SegmentationShusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li et al.ICCV 2021 · 124 citations
- MinVIS: A Minimal Video Instance Segmentation Framework without Video-based TrainingDe-An Huang, Zhiding Yu, Anima AnandkumarNeurIPS 2022 · 135 citations
- A Generalized Framework for Video Instance SegmentationMiran Heo, Sukjun Hwang, Jeongseok Hyun, Hanjung Kim et al.CVPR 2023
- MDQE: Mining Discriminative Query Embeddings to Segment Occluded Instances on Challenging VideosMinghan Li, Shuai Li, Wangmeng Xiang, Lei ZhangCVPR 2023
- TCOVIS: Temporally Consistent Online Video Instance SegmentationJunlong Li, Bingyao Yu, Yongming Rao, Jie Zhou et al.ICCV 2023 · 23 citations
