RELO: Reinforcement Learning to Localize for Visual Object Tracking
Xin Chen, Chuanyu Sun, Jiao Xu, Houwen Peng, Dong Wang, Huchuan Lu, Kede Ma
Abstract
Conventional visual object trackers localize targets using handcrafted spatial priors, often in the form of heatmaps. Such priors provide only surrogate supervision and are poorly aligned with tracking optimization and evaluation metrics, such as intersection over union (IoU) and area under the success curve (AUC). Here, we introduce RELO, a REinforcement-learning-to-LOcalize method for visual object tracking that formulates target localization as a Markov decision process. Specifically, RELO replaces handcrafted spatial priors with a localization policy learned over spatial positions via reinforcement learning, with rewards combining frame-level IoU and sequence-level AUC. We additionally introduce layer-aligned temporal token propagation to improve semantic consistency across frames, with negligible computational overhead. Across multiple benchmarks, RELO achieves superior results, attaining % AUC on LaSOT without template updates. This confirms that reward-driven localization provides an effective alternative to prior-driven localization for visual object tracking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97e25297-d438-4c3a-9fbd-693ca1906d7dBuilds on28
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation GuidelinesYinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan et al.AAAI 2020 · 944 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu et al.NeurIPS 2022 · 556 citations
Related papers
- Visual Tracking via Hierarchical Deep Reinforcement LearningDawei Zhang, Zhonglong Zheng, Riheng Jia, Minglu LiAAAI 2021 · 29 citations
- Autoregressive Visual TrackingXing Wei, Yifan Bai, Yongchao Zheng, Dahu Shi et al.CVPR 2023
- Online Decision Based Visual Tracking via Reinforcement LearningKe Song, Wei Zhang, Ran Song, Yibin LiNeurIPS 2020 · 20 citations
- AutoTrack: Towards High-Performance Visual Tracking for UAV With Automatic Spatio-Temporal RegularizationYiming Li, Changhong Fu, Fangqiang Ding, Ziyuan Huang et al.CVPR 2020
- STRONG: Spatio-Temporal Reinforcement Learning for Cross-Modal Video Moment LocalizationDa Cao, Yawen Zeng, Meng Liu, Xiangnan He et al.ACM MM 2020 · 47 citations
