Towards Open-Vocabulary Video Instance Segmentation
Haochen Wang, Xiaolong Jiang, Xu Tang, Yao Hu, Cilin Yan, Weidi Xie, Shuai Wang, Efstratios Gavves
摘要
Video Instance Segmentation (VIS) aims at segmenting and categorizing objects in videos from a closed set of training categories, lacking the generalization ability to handle novel categories in real-world videos. To address this limitation, we make the following three contributions. First, we introduce the novel task of Open-Vocabulary Video Instance Segmentation, which aims to simultaneously segment, track, and classify objects in videos from open-set categories, including novel categories unseen during training. Second, to benchmark Open-Vocabulary VIS, we collect a Large-Vocabulary Video Instance Segmentation dataset (LV-VIS), that contains well-annotated objects from 1,196 diverse categories, significantly surpassing the category size of existing datasets by more than one order of magnitude. Third, we propose an efficient Memory-Induced Transformer architecture, OV2Seg, to first achieve Open-Vocabulary VIS in an end-to-end manner with near real-time inference speed. Extensive experiments on LV-VIS and four existing VIS datasets demonstrate the strong zero-shot generalization ability of OV2Seg on novel categories. The dataset and code are released here https://github.com/haochenheheda/LVVIS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath 等ICLR 2026 · 被引用 1,103 次
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and GroundingChristopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park 等CVPR 2026 · 被引用 144 次
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and VideosWeifeng Lin, Xinyu Wei, Ruichuan An, Tianhe Ren 等NeurIPS 2025 · 被引用 47 次
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual ReasoningYe Liu, Zongyang Ma, Junfu Pu, Zhongang Qi 等NeurIPS 2025 · 被引用 39 次
- General Object Foundation Model for Images and Videos at ScaleJunfeng Wu, Yi Jiang, Qihao Liu, Zehuan Yuan 等CVPR 2024 · 被引用 36 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 被引用 615 次
相关 Paper
- Unidentified Video Objects: A Benchmark for Dense, Open-World SegmentationWeiyao Wang, Matt Feiszli, Heng Wang, Du TranICCV 2021 · 被引用 151 次
- OV-VOD: Open-Vocabulary Video Object DetectionZhihong Zheng, Yang Cao, Junlong Gao, Hanzi WangACM MM 2025
- OpenVIS: Open-vocabulary Video Instance SegmentationPinxue Guo, Hao Huang, Peiyang He, Xuefeng Liu 等AAAI 2025 · 被引用 26 次
- A Generalized Framework for Video Instance SegmentationMiran Heo, Sukjun Hwang, Jeongseok Hyun, Hanjung Kim 等CVPR 2023
- Open-Vocabulary Audio-Visual Semantic SegmentationRuohao Guo, Liao Qu, Dantong Niu, Yanyu Qi 等ACM MM 2024 · 被引用 4 次
