OpenVIS: Open-vocabulary Video Instance Segmentation
Pinxue Guo, Hao Huang, Peiyang He, Xuefeng Liu, Tianjun Xiao, Wenqiang Zhang
Abstract
Open-vocabulary Video Instance Segmentation (OpenVIS) can simultaneously detect, segment, and track arbitrary object categories in a video, without being constrained to categories seen during training. In this work, we propose In-stFormer, a carefully designed framework for the Open-VIS task that achieves powerful open-vocabulary capabilities through lightweight fine-tuning with limited-category data. InstFormer begins with the open-world mask proposal network, encouraged to propose all potential instance class-agnostic masks by the contrastive instance margin loss. Next, we introduce InstCLIP, adapted from pre-trained CLIP with Instance Guidance Attention, which encodes openvocabulary instance tokens efficiently. These instance tokens not only enable open-vocabulary classification but also offer strong universal tracking capabilities. Furthermore, to prevent the tracking module from being constrained by the training data with limited categories, we propose the universal rollout association, which transforms the tracking problem into predicting the next frame's instance tracking token. The experimental results demonstrate the proposed In-stFormer achieve state-of-the-art capabilities on a comprehensive OpenVIS evaluation benchmark, while also achieves competitive performance in fully supervised VIS task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 56885262-8549-4645-a62c-dd034d6ff8d1Cited by top-tier papers9
- X-Prompt: Multi-modal Visual Prompt for Video Object SegmentationPinxue Guo, Wanyun Li, Hao Huang, Lingyi Hong et al.ACM MM 2024 · 7 citations
- Sliced Wasserstein Bridge for Open-Vocabulary Video Instance SegmentationZheyun Qin, Deng Yu, Chuanchen Luo, Zhumin ChenICCV 2025 · 5 citations
- Open-Vocabulary Audio-Visual Semantic SegmentationRuohao Guo, Liao Qu, Dantong Niu, Yanyu Qi et al.ACM MM 2024 · 4 citations
- Aligning Instance Brownian Bridge with Texts for Open-Vocabulary Video Instance SegmentationZesen Cheng, Kehan Li, Hao Li, Peng Jin et al.AAAI 2025 · 3 citations
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal ObjectsGabriel Fiastre, Antoine Yang, Cordelia SchmidCVPR 2026 · 2 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- Open-Vocabulary Universal Image Segmentation with MaskCLIPZheng Ding, Jieke Wang, Zhuowen TuICML 2023 · 150 citations
- MasQCLIP for Open-Vocabulary Universal Image SegmentationXin Xu, Tianyi Xiong, Zheng Ding, Zhuowen TuICCV 2023 · 57 citations
- Towards Open-Vocabulary Video Instance SegmentationHaochen Wang, Xiaolong Jiang, Xu Tang, Yao Hu et al.ICCV 2023 · 56 citations
- Efficient Video Instance Segmentation via Tracklet Query and ProposalJialian Wu, Sudhir Yarram, Hui Liang, Tian Lan et al.CVPR 2022 · 33 citations
- ProposalCLIP: Unsupervised Open-Category Object Proposal Generation via Exploiting CLIP CuesHengcan Shi, Munawar Hayat, Yicheng Wu, Jianfei CaiCVPR 2022 · 59 citations
