CompFeat: Comprehensive Feature Aggregation for Video Instance Segmentation
Yang Fu, Linjie Yang, Ding Liu, Thomas S. Huang, Humphrey Shi
Abstract
Video instance segmentation is a complex task in which we need to detect, segment, and track each object for any given video. Previous approaches only utilize single-frame features for the detection, segmentation, and tracking of objects and they suffer in the video scenario due to several distinct challenges such as motion blur and drastic appearance change. To eliminate ambiguities introduced by only using singleframe features, we propose a novel comprehensive feature aggregation approach (CompFeat) to refine features at both frame-level and object-level with temporal and spatial context information. The aggregation process is carefully designed with a new attention mechanism which significantly increases the discriminative power of the learned features. We further improve the tracking capability of our model through a siamese design by incorporating both feature similarities and spatial similarities. Experiments conducted on the YouTube-VIS dataset validate the effectiveness of proposed CompFeat. Our code will be available at https://github.com/SHI-Labs/ CompFeat-for-Video-Instance-Segmentation .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Instances as QueriesYuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li et al.ICCV 2021 · 331 citations
- Video Instance Segmentation using Inter-Frame Communication TransformersSukjun Hwang, Miran Heo, Seoung Wug Oh, Seon Joo KimNeurIPS 2021 · 174 citations
- VITA: Video Instance Segmentation via Object Token AssociationMiran Heo, Sukjun Hwang, Seoung Wug Oh, Joon-Young Lee et al.NeurIPS 2022 · 146 citations
- Crossover Learning for Fast Online Video Instance SegmentationShusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li et al.ICCV 2021 · 124 citations
- Video K-Net: A Simple, Strong, and Unified Baseline for Video SegmentationXiangtai Li, Wenwei Zhang, Jiangmiao Pang, Kai Chen et al.CVPR 2022 · 71 citations
Builds on2
Related papers
- Classifying, Segmenting, and Tracking Object Instances in Video with Mask PropagationGedas Bertasius, Lorenzo TorresaniCVPR 2020
- End-to-End Video Instance Segmentation via Spatial-Temporal Graph Neural NetworksTao Wang, Ning Xu, Kean Chen, Weiyao LinICCV 2021 · 30 citations
- Prototypical Cross-Attention Networks for Multiple Object Tracking and SegmentationLei Ke, Xia Li, Martin Danelljan, Yu-Wing Tai et al.NeurIPS 2021 · 92 citations
- Spatial Feature Calibration and Temporal Fusion for Effective One-Stage Video Instance SegmentationMinghan Li, Shuai Li, Lida Li, Lei ZhangCVPR 2021
- SG-Net: Spatial Granularity Network for One-Stage Video Instance SegmentationDongfang Liu, Yiming Cui, Wenbo Tan, Yingjie Victor ChenCVPR 2021
