Video OWL-ViT: Temporally-consistent open-world localization in video
Georg Heigold, Daniel Keysers, Matthias Minderer, Mario Lucic, Alexey A. Gritsenko, Fisher Yu, Alex Bewley, Thomas Kipf
Abstract
We present an architecture and a training recipe that adapts pre-trained open-world image models to localization in videos. Understanding the open visual world (without being constrained by fixed label spaces) is crucial for many real-world vision tasks. Contrastive pre-training on large image-text datasets has recently led to significant improvements for image-level tasks. For more structured tasks involving object localization applying pre-trained models is more challenging. This is particularly true for video tasks, where task-specific data is limited. We show successful transfer of open-world models by building on the OWL-ViT open-vocabulary detection model and adapting it to video by adding a transformer decoder. The decoder propagates object representations recurrently through time by using the output tokens for one frame as the object queries for the next. Our model is end-to-end trainable on video data and enjoys improved temporal consistency compared to tracking-by-detection baselines, while retaining the open-world capabilities of the backbone detector. We evaluate our model on the challenging TAO-OW benchmark and demonstrate that open-world capabilities, learned from large-scale image-text pre-training, can be transferred successfully to open-world localization across diverse videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Moving Off-the-Grid: Scene-Grounded Video RepresentationsSjoerd van Steenkiste, Daniel Zoran, Yi Yang, Yulia Rubanova et al.NeurIPS 2024 · 13 citations
- Naver: a Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic ReasoningZhixi Cai, Fucai Ke, Simindokht Jahangard, Maria Garcia de la Banda et al.ICCV 2025 · 1 citation
- NetTrack: Tracking Highly Dynamic Objects with a NetGuangze Zheng, Shijie Lin, Haobo Zuo, Changhong Fu et al.CVPR 2024
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
Related papers
- Contrastive Feature Masking Open-Vocabulary Vision TransformerDahun Kim, Anelia Angelova, Weicheng KuoICCV 2023 · 40 citations
- Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision TransformersDahun Kim, Anelia Angelova, Weicheng KuoCVPR 2023
- Scaling Open-Vocabulary Object DetectionMatthias Minderer, Alexey A. Gritsenko, Neil HoulsbyNeurIPS 2023 · 482 citations
- QDETRv: Query-Guided DETR for One-Shot Object Localization in VideosYogesh Kumar, Saswat Mallick, Anand Mishra, Sowmya Rasipuram et al.AAAI 2024 · 4 citations
- CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense PredictionSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin et al.ICLR 2024 · 129 citations
