OVTrack: Open-Vocabulary Multiple Object Tracking
Siyuan Li, Tobias Fischer, Lei Ke, Henghui Ding, Martin Danelljan, Fisher Yu
摘要
The ability to recognize, localize and track dynamic objects in a scene is fundamental to many real-world applications, such as self-driving and robotic systems. Yet, traditional multiple object tracking (MOT) benchmarks rely only on a few object categories that hardly represent the multitude of possible objects that are encountered in the real world. This leaves contemporary MOT methods limited to a small set of pre-defined object categories. In this paper, we address this limitation by tackling a novel task, openvocabulary MOT, that aims to evaluate tracking beyond predefined training categories. We further develop OVTrack, an open-vocabulary tracker that is capable of tracking arbitrary object classes. Its design is based on two key ingredients: First, leveraging vision-language models for both classification and association via knowledge distillation; second, a data hallucination strategy for robust appearance feature learning from denoising diffusion probabilistic models. The result is an extremely data-efficient open-vocabulary tracker that sets a new state-of-the-art on the large-scale, largevocabulary TAO benchmark, while being trained solely on static images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Type-to-Track: Retrieve Any Object via Prompt-based TrackingPha A. Nguyen, Kha Gia Quach, Kris Kitani, Khoa LuuNeurIPS 2023 · 被引用 38 次
- General Object Foundation Model for Images and Videos at ScaleJunfeng Wu, Yi Jiang, Qihao Liu, Zehuan Yuan 等CVPR 2024 · 被引用 36 次
- Video OWL-ViT: Temporally-consistent open-world localization in videoGeorg Heigold, Daniel Keysers, Matthias Minderer, Mario Lucic 等ICCV 2023 · 被引用 22 次
- Unsupervised Video Deraining with An Event CameraJin Wang, Wenming Weng, Yueyi Zhang, Zhiwei XiongICCV 2023 · 被引用 21 次
- Self-Supervised Multi-Object Tracking with Path ConsistencyZijia Lu, Bing Shuai, Yanbei Chen, Zhenlin Xu 等CVPR 2024 · 被引用 13 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- GLATrack: Global and Local Awareness for Open-Vocabulary Multiple Object TrackingGuangyao Li, Yajun Jian, Yan Yan, Hanzi WangACM MM 2024 · 被引用 3 次
- VOVTrack: Exploring the Potentiality in Raw Videos for Open-Vocabulary Multi-Object TrackingZekun Qian, Ruize Han, Junhui Hou, Linqi Song 等ICCV 2025 · 被引用 3 次
- COVTrack: Continuous Open-Vocabulary Tracking via Adaptive Multi-Cue FusionZekun Qian, Ruize Han, Zhixiang Wang, Junhui Hou 等ICCV 2025 · 被引用 1 次
- Attention to Trajectory: Trajectory-Aware Open-Vocabulary TrackingYunhao Li, Yifan Jiao, Dan Meng, Heng Fan 等ICCV 2025 · 被引用 1 次
- DOVTrack: Data-Efficient Open-Vocabulary TrackingZekun Qian, Ruize Han, Zhixiang Wang, Junhui Hou 等NeurIPS 2025 · 被引用 1 次
