Prototypical Cross-Attention Networks for Multiple Object Tracking and Segmentation
Lei Ke, Xia Li, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu
Abstract
Multiple object tracking and segmentation requires detecting, tracking, and segmenting objects belonging to a set of given classes. Most approaches only exploit the temporal dimension to address the association problem, while relying on single frame predictions for the segmentation mask itself. We propose Prototypical Cross-Attention Network (PCAN), capable of leveraging rich spatio-temporal information for online multiple object tracking and segmentation. PCAN first distills a space-time memory into a set of prototypes and then employs cross-attention to retrieve rich information from the past frames. To segment each object, PCAN adopts a prototypical appearance module to learn a set of contrastive foreground and background prototypes, which are then propagated over time. Extensive experiments demonstrate that PCAN outperforms current video instance tracking and segmentation competition winners on both Youtube-VIS and BDD100K datasets, and shows efficacy to both one-stage and two-stage segmentation frameworks. Code and video resources are available at http://vis.xyz/pub/pcan.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5da2f69-58c0-4300-b793-521c8ec54aa1Cited by top-tier papers18
- VITA: Video Instance Segmentation via Object Token AssociationMiran Heo, Sukjun Hwang, Seoung Wug Oh, Joon-Young Lee et al.NeurIPS 2022 · 146 citations
- Learning Cross-Image Object Semantic Relation in Transformer for Few-Shot Fine-Grained Image ClassificationBo Zhang, Jiakang Yuan, Baopu Li, Tao Chen et al.ACM MM 2022 · 42 citations
- TCOVIS: Temporally Consistent Online Video Instance SegmentationJunlong Li, Bingyao Yu, Yongming Rao, Jie Zhou et al.ICCV 2023 · 23 citations
- InsPro: Propagating Instance Query and Proposal for Online Video Instance SegmentationFei He, Haoyang Zhang, Naiyu Gao, Jian Jia et al.NeurIPS 2022 · 23 citations
- Video Task Decathlon: Unifying Image and Video Tasks in Autonomous DrivingThomas E. Huang, Yifan Liu, Luc Van Gool, Fisher YuICCV 2023 · 10 citations
Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
Related papers
- Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object SegmentationTianfei Zhou, Jianwu Li, Xueyi Li, Ling ShaoCVPR 2021
- CompFeat: Comprehensive Feature Aggregation for Video Instance SegmentationYang Fu, Linjie Yang, Ding Liu, Thomas S. Huang et al.AAAI 2021 · 77 citations
- Context-Aware Relative Object Queries to Unify Video Instance and Panoptic SegmentationAnwesa Choudhuri, Girish Chowdhary, Alexander G. SchwingCVPR 2023
- Classifying, Segmenting, and Tracking Object Instances in Video with Mask PropagationGedas Bertasius, Lorenzo TorresaniCVPR 2020
- Learning Spatial-Semantic Features for Robust Video Object SegmentationXin Li, Deshui Miao, Zhenyu He, Yaowei Wang et al.ICLR 2025
