Reading Relevant Feature from Global Representation Memory for Visual Object Tracking
Xinyu Zhou, Pinxue Guo, Lingyi Hong, Jinglun Li, Wei Zhang, Weifeng Ge, Wenqiang Zhang
Abstract
Reference features from a template or historical frames are crucial for visual object tracking. Prior works utilize all features from a fixed template or memory for visual object tracking. However, due to the dynamic nature of videos, the required reference historical information for different search regions at different time steps is also inconsistent. Therefore, using all features in the template and memory can lead to redundancy and impair tracking performance. To alleviate this issue, we propose a novel tracking paradigm, consisting of a relevance attention mechanism and a global representation memory, which can adaptively assist the search region in selecting the most relevant historical information from reference features. Specifically, the proposed relevance attention mechanism in this work differs from previous approaches in that it can dynamically choose and build the optimal global representation memory for the current frame by accessing cross-frame information globally. Moreover, it can flexibly read the relevant historical information from the constructed memory to reduce redundancy and counteract the negative effects of harmful information. Extensive experiments validate the effectiveness of the proposed method, achieving competitive performance on five challenging datasets with 71 FPS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ddffd22-702a-443a-a3f0-3cd4027d08c0Cited by top-tier papers15
- ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelYiming Sun, Fan Yu, Shaoxiang Chen, Yu Zhang et al.NeurIPS 2024 · 21 citations
- DeTrack: In-model Latent Denoising Learning for Visual Object TrackingXinyu Zhou, Jinglun Li, Lingyi Hong, Kaixun Jiang et al.NeurIPS 2024 · 14 citations
- X-Prompt: Multi-modal Visual Prompt for Video Object SegmentationPinxue Guo, Wanyun Li, Hao Huang, Lingyi Hong et al.ACM MM 2024 · 7 citations
- General Compression Framework for Efficient Transformer Object TrackingLingyi Hong, Jinglun Li, Xinyu Zhou, Shilin Yan et al.ICCV 2025 · 5 citations
- DINTR: Tracking via Diffusion-based InterpolationPha A. Nguyen, Ngan Le, Jackson David Cothren, Alper Yilmaz et al.NeurIPS 2024 · 5 citations
Builds on26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
Related papers
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang et al.ICCV 2019 · 109 citations
- Less Is More: Token Context-Aware Learning for Object TrackingChenlong Xu, Bineng Zhong, Qihua Liang, Yaozong Zheng et al.AAAI 2025
- Scoring, Remember, and Reference: Catching Camouflaged Objects in VideosYu'ang Feng, Shuyong Gao, Fuzhen Yan, Yicheng Song et al.ICCV 2025 · 2 citations
- Robust Object Modeling for Visual TrackingYidong Cai, Jie Liu, Jie Tang, Gangshan WuICCV 2023 · 165 citations
- Explicit Context Reasoning with Supervision for Visual TrackingFansheng Zeng, Bineng Zhong, Haiying Xia, Yufei Tan et al.ACM MM 2025 · 1 citation
