InstanceFormer: An Online Video Instance Segmentation Framework
Rajat Koner, Tanveer Hannan, Suprosanna Shit, Sahand Sharifzadeh, Matthias Schubert, Thomas Seidl, Volker Tresp
Abstract
Recent transformer-based offline video instance segmentation (VIS) approaches achieve encouraging results and significantly outperform online approaches. However, their reliance on the whole video and the immense computational complexity caused by full Spatio-temporal attention limit them in real-life applications such as processing lengthy videos. In this paper, we propose a single-stage transformer-based efficient online VIS framework named InstanceFormer, which is especially suitable for long and challenging videos. We propose three novel components to model short-term and long-term dependency and temporal coherence. First, we propagate the representation, location, and semantic information of prior instances to model short-term changes. Second, we propose a novel memory cross-attention in the decoder, which allows the network to look into earlier instances within a certain temporal window. Finally, we employ a temporal contrastive loss to impose coherence in the representation of an instance across all frames. Memory attention and temporal coherence are particularly beneficial to long-range dependency modeling, including challenging scenarios like occlusion. The proposed InstanceFormer outperforms previous online benchmark methods by a large margin across multiple datasets. Most importantly, InstanceFormer surpasses offline approaches for challenging and long datasets such as YouTube-VIS-2021 and OVIS. Code is available at https://github.com/rajatkoner08/InstanceFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Temporal Collection and Distribution for Referring Video Object SegmentationJiajin Tang, Ge Zheng, Sibei YangICCV 2023 · 44 citations
- MobileInst: Video Instance Segmentation on the MobileRenhong Zhang, Tianheng Cheng, Shusheng Yang, Haoyi Jiang et al.AAAI 2024 · 10 citations
- OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and CaptioningAnwesa Choudhuri, Girish Chowdhary, Alexander G. SchwingNeurIPS 2024 · 7 citations
- Clustering Propagation for Universal Medical Image SegmentationYuhang Ding, Liulei Li, Wenguan Wang, Yi YangCVPR 2024
- MDQE: Mining Discriminative Query Embeddings to Segment Occluded Instances on Challenging VideosMinghan Li, Shuai Li, Wangmeng Xiang, Lei ZhangCVPR 2023
Builds on16
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Instances as QueriesYuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li et al.ICCV 2021 · 331 citations
Related papers
- VITA: Video Instance Segmentation via Object Token AssociationMiran Heo, Sukjun Hwang, Seoung Wug Oh, Joon-Young Lee et al.NeurIPS 2022 · 146 citations
- Hybrid Instance-Aware Temporal Fusion for Online Video Instance SegmentationXiang Li, Jinglu Wang, Xiao Li, Yan LuAAAI 2022 · 20 citations
- Temporally Efficient Vision Transformer for Video Instance SegmentationShusheng Yang, Xinggang Wang, Yu Li, Yuxin Fang et al.CVPR 2022 · 68 citations
- Video Instance Segmentation using Inter-Frame Communication TransformersSukjun Hwang, Miran Heo, Seoung Wug Oh, Seon Joo KimNeurIPS 2021 · 174 citations
- A Generalized Framework for Video Instance SegmentationMiran Heo, Sukjun Hwang, Jeongseok Hyun, Hanjung Kim et al.CVPR 2023
