SAM2-OV: A Novel Detection-Only Tuning Paradigm for Open-Vocabulary Multi-Object Tracking
Yangkai Chen, Qiangqiang Wu, Guangyao Li, Junlong Gao, Guanglin Niu, Hanzi Wang
摘要
Open-vocabulary multi-object tracking (OV-MOT) aims to track objects with unseen categories beyond the training set. While existing methods rely on pseudo video sequences synthesized from static images, they struggle to model realistic motion patterns, resulting in limited association performance in real-world scenarios. To alleviate these issues, we propose SAM2-OV, a novel association learning-free OV-MOT method that adopts a detection-only tuning paradigm, eliminating the need for synthetic sequences or spatiotemporal supervision and substantially reducing the overall learnable parameters. The core of our method is a Unified Detection Module (UDM), which effectively provides object-level prompts to enable SAM2 for OV-MOT. Enabled by UDM, SAM2-OV is the first to integrate SAM2 for OV-MOT, fully unleashing its zero-shot cross-frame association ability. To further enhance object association under occlusion and abrupt motion, we introduce a Motion Prior Assistance Module (MPAM) that incorporates motion cues into the mask selection process. In addition, a Semantic Enhancement Adapter (SEA) distilled from CLIP is used to improve classification generalization. A sparse prompting strategy is also adopted to reduce computational redundancy by triggering detection only on selected keyframes. As only the detection module is tuned on static images, the overall training process remains simple and efficient. Experiments on the TAO dataset demonstrate that SAM2-OV achieves state-of-the-art performance under the TETA metric, particularly on novel categories. Evaluations on the KITTI dataset show the strong zero-shot cross-domain transferability of our SAM2-OV.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 被引用 927 次
- Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelYu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi 等CVPR 2022 · 被引用 311 次
- MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object TrackingRuopeng Gao, Limin WangICCV 2023 · 被引用 143 次
相关 Paper
- SAM2MOT: A Novel Paradigm of Multi-Object Tracking by SegmentationJunjie Jiang, Zelin Wang, Manqi Zhao, Yin Li 等AAAI 2026 · 被引用 19 次
- Matching Anything by Segmenting AnythingSiyuan Li, Lei Ke, Martin Danelljan, Luigi Piccinelli 等CVPR 2024
- VOVTrack: Exploring the Potentiality in Raw Videos for Open-Vocabulary Multi-Object TrackingZekun Qian, Ruize Han, Junhui Hou, Linqi Song 等ICCV 2025 · 被引用 3 次
- Attention to Trajectory: Trajectory-Aware Open-Vocabulary TrackingYunhao Li, Yifan Jiao, Dan Meng, Heng Fan 等ICCV 2025 · 被引用 1 次
- A Simple Baseline for Open-World Tracking via Self-trainingBingyang Wang, Tanlin Li, Jiannan Wu, Yi Jiang 等ACM MM 2023 · 被引用 1 次
