DOVTrack: Data-Efficient Open-Vocabulary Tracking
Zekun Qian, Ruize Han, Zhixiang Wang, Junhui Hou, Wei Feng
摘要
Open-Vocabulary Multi-Object Tracking (OVMOT) aims to detect and track multicategory objects including both seen and unseen categories during training. Currently, a significant challenge in this domain is the lack of large-scale annotated video data for training. To address this challenge, this work aims to effectively train the OV tracker using only the existing limited and sparsely annotated video data. We propose a comprehensive training sample space expansion strategy that addresses the fundamental limitation of sparse annotations in OVMOT training. Specifically, for the association task, we develop a diffusion-based feature generation framework that synthesizes intermediate object features between sparsely annotated frames, effectively expanding the training sample space by approximately 3× and enabling robust association learning from temporally continuous features. For the detection task, we introduce a dynamic group contrastive learning approach that generates diverse sample groups through affinity, dispersion, and adversarial grouping strategies, tripling the effective training samples for classification while maintaining sample quality. Additionally, we propose an adaptive localization loss that expands positive sample coverage by lowering IoU thresholds while mitigating noise through confidence-based weighting. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the OVMOT benchmark, surpassing existing methods by 3.8% in TETA metric, without requiring additional data or annotations. The code will be available at https://github.com/zekunqian/DOVTrack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object DetectionYupeng Zhang, Ruize Han, Zhiwei Chen, Wei Feng 等CVPR 2026 · 被引用 2 次
- InfoScan: Information-Efficient Visual Scanning via Resource-Adaptive WalksYifeng Wu, Huimin Huang, Shangjie Zhou, Yawen Huang 等ICLR 2026
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 被引用 1,030 次
相关 Paper
- COVTrack: Continuous Open-Vocabulary Tracking via Adaptive Multi-Cue FusionZekun Qian, Ruize Han, Zhixiang Wang, Junhui Hou 等ICCV 2025 · 被引用 1 次
- VOVTrack: Exploring the Potentiality in Raw Videos for Open-Vocabulary Multi-Object TrackingZekun Qian, Ruize Han, Junhui Hou, Linqi Song 等ICCV 2025 · 被引用 3 次
- OVTrack: Open-Vocabulary Multiple Object TrackingSiyuan Li, Tobias Fischer, Lei Ke, Henghui Ding 等CVPR 2023
- GLATrack: Global and Local Awareness for Open-Vocabulary Multiple Object TrackingGuangyao Li, Yajun Jian, Yan Yan, Hanzi WangACM MM 2024 · 被引用 3 次
- SAM2-OV: A Novel Detection-Only Tuning Paradigm for Open-Vocabulary Multi-Object TrackingYangkai Chen, Qiangqiang Wu, Guangyao Li, Junlong Gao 等AAAI 2026
