Video Instance Segmentation using Inter-Frame Communication Transformers
Sukjun Hwang, Miran Heo, Seoung Wug Oh, Seon Joo Kim
Abstract
We propose a novel end-to-end solution for video instance segmentation (VIS) based on transformers. Recently, the per-clip pipeline shows superior performance over per-frame methods leveraging richer information from multiple frames. However, previous per-clip models require heavy computation and memory usage to achieve frame-to-frame communications, limiting practicality. In this work, we propose Inter-frame Communication Transformers (IFC), which significantly reduces the overhead for information-passing between frames by efficiently encoding the context within the input clip. Specifically, we propose to utilize concise memory tokens as a mean of conveying information as well as summarizing each frame scene. The features of each frame are enriched and correlated with other frames through exchange of information between the precisely encoded memory tokens. We validate our method on the latest benchmark sets and achieved the state-of-the-art performance (AP 44.6 on YouTube-VIS 2019 val set using the offline inference) while having a considerably fast runtime (89.4 FPS). Our method can also be applied to near-online inference for processing a video in real-time with only a small delay. The code will be made available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f4a50d9-409b-4616-a66e-c6f0c5cba28eCited by top-tier papers13
- Language as Queries for Referring Video Object SegmentationJiannan Wu, Yi Jiang, Peize Sun, Zehuan Yuan et al.CVPR 2022 · 143 citations
- CTVIS: Consistent Training for Online Video Instance SegmentationKaining Ying, Qing Zhong, Weian Mao, Zhenhua Wang et al.ICCV 2023 · 72 citations
- Temporally Efficient Vision Transformer for Video Instance SegmentationShusheng Yang, Xinggang Wang, Yu Li, Yuxin Fang et al.CVPR 2022 · 68 citations
- Large-scale Video Panoptic Segmentation in the Wild: A BenchmarkJiaxu Miao, Xiaohan Wang, Yu Wu, Wei Li et al.CVPR 2022 · 58 citations
- InsPro: Propagating Instance Query and Proposal for Online Video Instance SegmentationFei He, Haoyang Zhang, Naiyu Gao, Jian Jia et al.NeurIPS 2022 · 23 citations
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Crossover Learning for Fast Online Video Instance SegmentationShusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li et al.ICCV 2021 · 124 citations
Related papers
- End-to-End Video Instance Segmentation With TransformersYuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen et al.CVPR 2021
- VITA: Video Instance Segmentation via Object Token AssociationMiran Heo, Sukjun Hwang, Seoung Wug Oh, Joon-Young Lee et al.NeurIPS 2022 · 146 citations
- Efficient Video Instance Segmentation via Tracklet Query and ProposalJialian Wu, Sudhir Yarram, Hui Liang, Tian Lan et al.CVPR 2022 · 33 citations
- InstanceFormer: An Online Video Instance Segmentation FrameworkRajat Koner, Tanveer Hannan, Suprosanna Shit, Sahand Sharifzadeh et al.AAAI 2023 · 16 citations
- SyncVIS: Synchronized Video Instance SegmentationRongkun Zheng, Lu Qi, Xi Chen, Yi Wang et al.NeurIPS 2024 · 8 citations
