Temporal-Aware Query Routing for Real-Time Video Instance Segmentation
Zesen Cheng, Kehan Li, Yian Zhao, Hang Zhang, Chang Liu, Jie Chen
Abstract
With the rise of applications such as embodied intelligence, developing high real-time online video instance segmentation (VIS) has become increasingly important. However, through time profiling of the components in advanced online VIS architecture (i.e., transformer-based architecture), we find that the transformer decoder significantly hampers the inference speed. Further analysis of the similarities between the outputs from adjacent frames at each transformer decoder layer reveals significant redundant computations within the transformer decoder. To address this issue, we introduce Temporal-Aware query Routing (TAR) mechanism. We embed it before each transformer decoder layer. By fusing the optimal queries from the previous frame, the queries output by the preceding decoder layer, and their differential information, TAR predicts a binary classification score and then uses an argmax operation to determine whether the current layer should be skipped. Experimental results demonstrate that integrating TAR into the baselines achieves significant efficiency gains (24.7
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08b1fc85-59d9-4b03-86a5-6f267812c229Builds on25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
Related papers
- Temporally Efficient Vision Transformer for Video Instance SegmentationShusheng Yang, Xinggang Wang, Yu Li, Yuxin Fang et al.CVPR 2022 · 68 citations
- Video Semantic Segmentation via Sparse Temporal TransformerJiangtong Li, Wentao Wang, Junjie Chen, Li Niu et al.ACM MM 2021 · 47 citations
- Video Instance Segmentation using Inter-Frame Communication TransformersSukjun Hwang, Miran Heo, Seoung Wug Oh, Seon Joo KimNeurIPS 2021 · 174 citations
- End-to-End Video Instance Segmentation With TransformersYuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen et al.CVPR 2021
- Efficient Video Instance Segmentation via Tracklet Query and ProposalJialian Wu, Sudhir Yarram, Hui Liang, Tian Lan et al.CVPR 2022 · 33 citations
