Context-Aware Relative Object Queries to Unify Video Instance and Panoptic Segmentation
Anwesa Choudhuri, Girish Chowdhary, Alexander G. Schwing
Abstract
Object queries have emerged as a powerful abstraction to generically represent object proposals. However, their use for temporal tasks like video segmentation poses two questions: 1) How to process frames sequentially and propagate object queries seamlessly across frames. Using independent object queries per frame doesn't permit tracking, and requires post-processing. 2) How to produce temporally consistent, yet expressive object queries that model both appearance and position changes. Using the entire video at once doesn't capture position changes and doesn't scale to long videos. As one answer to both questions we propose 'context-aware relative object queries', which are continuously propagated frame-by-frame. They seamlessly track objects and deal with occlusion and re-appearance of objects, without post-processing. Further, we find contextaware relative object queries better capture position changes of objects in motion. We evaluate the proposed approach across three challenging tasks: video instance segmentation, multi-object tracking and segmentation, and video panoptic segmentation. Using the same approach and architecture, we match or surpass state-of-the art results on the diverse and challenging OVIS, Youtube-VIS, Cityscapes-VPS, MOTS 2020 and KITTI-MOTS data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 685db017-c326-4058-abda-7a881564168fCited by top-tier papers5
- Tracking Anything with Decoupled Video SegmentationHo Kei Cheng, Seoung Wug Oh, Brian L. Price, Alexander G. Schwing et al.ICCV 2023 · 240 citations
- OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and CaptioningAnwesa Choudhuri, Girish Chowdhary, Alexander G. SchwingNeurIPS 2024 · 7 citations
- CAVIS: Context-Aware Video Instance SegmentationSeunghun Lee, Jiwan Seo, Kiljoon Han, Minwoo Choi et al.ICCV 2025 · 4 citations
- Scene-Centric Unsupervised Video Panoptic SegmentationChristoph Reich, Oliver Hahn, Nikita Araslanov, Laura Leal-Taixe et al.CVPR 2026 · 1 citation
- Hierarchical Visual Prompt Learning for Continual Video Instance SegmentationJiahua Dong, Hui Yin, Wenqi Liang, Hanbin Zhao et al.ICCV 2025
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
Related papers
- Learning Spatial-Semantic Features for Robust Video Object SegmentationXin Li, Deshui Miao, Zhenyu He, Yaowei Wang et al.ICLR 2025
- InsPro: Propagating Instance Query and Proposal for Online Video Instance SegmentationFei He, Haoyang Zhang, Naiyu Gao, Jian Jia et al.NeurIPS 2022 · 23 citations
- Classifying, Segmenting, and Tracking Object Instances in Video with Mask PropagationGedas Bertasius, Lorenzo TorresaniCVPR 2020
- CompFeat: Comprehensive Feature Aggregation for Video Instance SegmentationYang Fu, Linjie Yang, Ding Liu, Thomas S. Huang et al.AAAI 2021 · 77 citations
- TarViS: A Unified Approach for Target-Based Video SegmentationAli Athar, Alexander Hermans, Jonathon Luiten, Deva Ramanan et al.CVPR 2023
