ODTrack: Online Dense Temporal Token Learning for Visual Tracking
Yaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo, Shengping Zhang, Xianxian Li
摘要
Online contextual reasoning and association across consecutive video frames are critical to perceive instances in visual tracking. However, most current top-performing trackers persistently lean on sparse temporal relationships between reference and search frames via an offline mode. Consequently, they can only interact independently within each image-pair and establish limited temporal correlations. To alleviate the above problem, we propose a simple, flexible and effective video-level tracking pipeline, named ODTrack, which densely associates the contextual relationships of video frames in an online token propagation manner. ODTrack receives video frames of arbitrary length to capture the spatio-temporal trajectory relationships of an instance, and compresses the discrimination features (localization information) of a target into a token sequence to achieve frame-to-frame association. This new solution brings the following benefits: 1) the purified token sequences can serve as prompts for the inference in the next video frame, whereby past information is leveraged to guide future inference; 2) the complex online update strategies are effectively avoided by the iterative propagation of token sequences, and thus we can achieve more efficient model representation and computation. ODTrack achieves a new SOTA performance on seven benchmarks, while running at real-time speed. Code and models are available at https://github.com/GXNU-ZhongLab/ODTrack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Exploring Enhanced Contextual Information for Video-Level Object TrackingBen Kang, Xin Chen, Simiao Lai, Yang Liu 等AAAI 2025 · 被引用 48 次
- Decoupled Spatio-Temporal Consistency Learning for Self-Supervised TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Ning Li 等AAAI 2025 · 被引用 41 次
- Robust Tracking via Mamba-based Context-aware Token LearningJinxia Xie, Bineng Zhong, Qihua Liang, Ning Li 等AAAI 2025 · 被引用 36 次
- SAM2LONG: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory TreeShuangrui Ding, Rui Qian, Xiaoyi Dong, Pan Zhang 等ICCV 2025 · 被引用 15 次
- SUTrack: Towards Simple and Unified Single Object TrackingXin Chen, Ben Kang, Wanting Geng, Jiawen Zhu 等AAAI 2025 · 被引用 12 次
它引用的顶会 Paper19
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 被引用 927 次
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 被引用 746 次
- Learning the Model Update for Siamese TrackersLichao Zhang, Abel Gonzalez-Garcia, Joost van de Weijer, Martin Danelljan 等ICCV 2019 · 被引用 371 次
- TCTrack: Temporal Contexts for Aerial TrackingZiang Cao, Ziyuan Huang, Liang Pan, Shiwei Zhang 等CVPR 2022 · 被引用 233 次
相关 Paper
- Less Is More: Token Context-Aware Learning for Object TrackingChenlong Xu, Bineng Zhong, Qihua Liang, Yaozong Zheng 等AAAI 2025
- Explicit Visual Prompts for Visual Object TrackingLiangtao Shi, Bineng Zhong, Qihua Liang, Ning Li 等AAAI 2024 · 被引用 117 次
- Explicit Context Reasoning with Supervision for Visual TrackingFansheng Zeng, Bineng Zhong, Haiying Xia, Yufei Tan 等ACM MM 2025 · 被引用 1 次
- Context-Aware Relative Object Queries to Unify Video Instance and Panoptic SegmentationAnwesa Choudhuri, Girish Chowdhary, Alexander G. SchwingCVPR 2023
- VideoTrack: Learning to Track Objects via Video TransformerFei Xie, Lei Chu, Jiahao Li, Yan Lu 等CVPR 2023
