TarViS: A Unified Approach for Target-Based Video Segmentation
Ali Athar, Alexander Hermans, Jonathon Luiten, Deva Ramanan, Bastian Leibe
Abstract
The general domain of video segmentation is currently fragmented into different tasks spanning multiple benchmarks. Despite rapid progress in the state-of-the-art, current methods are overwhelmingly task-specific and cannot conceptually generalize to other tasks. Inspired by recent approaches with multi-task capability, we propose TarViS: a novel, unified network architecture that can be applied to any task that requires segmenting a set of arbitrarily defined 'targets' in video. Our approach is flexible with respect to how tasks define these targets, since it models the latter as abstract 'queries' which are then used to predict pixel-precise target masks. A single TarViS model can be trained jointly on a collection of datasets spanning different tasks, and can hot-swap between tasks during inference without any task-specific retraining. To demonstrate its effectiveness, we apply TarViS to four different tasks, namely Video Instance Segmentation (VIS), Video Panoptic Segmentation (VPS), Video Object Segmentation (VOS) and Point Exemplar-guided Tracking (PET). Our unified, jointly trained model achieves state-of-the-art performance on 5/7 benchmarks spanning these four tasks, and competitive performance on the remaining two. Code and model weights are available at: https://github.com/Ali2500/TarViS
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73f2ca11-8871-45cb-8ce1-eb6a889fc6d9Cited by top-tier papers16
- Tracking Anything with Decoupled Video SegmentationHo Kei Cheng, Seoung Wug Oh, Brian L. Price, Alexander G. Schwing et al.ICCV 2023 · 240 citations
- DVIS: Decoupled Video Instance Segmentation FrameworkTao Zhang, Xingye Tian, Yu Wu, Shunping Ji et al.ICCV 2023 · 86 citations
- Segment Every Reference Object in Spatial and Temporal SpacesJiannan Wu, Yi Jiang, Bin Yan, Huchuan Lu et al.ICCV 2023 · 29 citations
- RMem: Restricted Memory Banks Improve Video Object SegmentationJunbao Zhou, Ziqi Pang, Yu-Xiong WangCVPR 2024 · 18 citations
- Advancing Complex Video Object Segmentation via Progressive Concept ConstructionZhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He et al.ICLR 2026 · 17 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- UniVS: Unified and Universal Video Segmentation with Prompts as QueriesMinghan Li, Shuai Li, Xindong Zhang, Lei ZhangCVPR 2024
- OMG-Seg: Is One Model Good Enough for all Segmentation?Xiangtai Li, Haobo Yuan, Wei Li, Henghui Ding et al.CVPR 2024
- A Unified Sequence Interface for Vision TasksTing Chen, Saurabh Saxena, Lala Li, Tsung-Yi Lin et al.NeurIPS 2022 · 201 citations
- Video Panoptic SegmentationDahun Kim, Sanghyun Woo, Joon-Young Lee, In So KweonCVPR 2020
- TubeFormer-DeepLab: Video Mask TransformerDahun Kim, Jun Xie, Huiyu Wang, Siyuan Qiao et al.CVPR 2022 · 28 citations
