LoRATv2: Enabling Low-Cost Temporal Modeling in One-Stream Trackers
Liting Lin, Heng Fan, Zhipeng Zhang, Yuqing Huang, Yaowei Wang, Yong Xu, Haibin Ling
Abstract
Transformer-based algorithms, such as LoRAT, have significantly enhanced objecttracking performance. However, these approaches rely on a standard attention mechanism, which incurs quadratic token complexity, making real-time inference computationally expensive. In this paper, we introduce LoRATv2, a novel tracking framework that addresses these limitations with three main contributions. First, LoRATv2 integrates frame-wise causal attention, which ensures full selfattention within each frame while enabling causal dependencies across frames, significantly reducing computational overhead. Moreover, key-value (KV) caching is employed to efficiently reuse past embeddings for further speedup. Second, building on LoRAT's parameter-efficient fine-tuning, we propose Stream-Specific LoRA Adapters (SSLA). As frame-wise causal attention introduces asymmetry in how streams access temporal information, SSLA assigns dedicated LoRA modules to the template and each search stream, with the main ViT backbone remaining frozen. This allows specialized adaptation for each stream's role in temporal tracking. Third, we introduce a two-phase progressive training strategy, which first trains a single-search-frame tracker and then gradually extends it to multi-searchframe inputs by introducing additional LoRA modules. This curriculum-based learning paradigm improves long-term tracking while maintaining training efficiency. In extensive experiments on multiple benchmarks, LoRATv2 achieves state-of-the-art performance, substantially improved efficiency, and a superior performance-to-FLOPs ratio over state-of-the-art trackers. The code is available at https://github.com/LitingLin/LoRATv2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34a2ad0f-938f-4618-9062-525f1f3bb2d7Cited by top-tier papers2
- Drift-Resilient Temporal Priors for Visual TrackingYuqing Huang, Liting Lin, Weijun Zhuang, Zhenyu He et al.CVPR 2026 · 1 citation
- Interactive Tracking: A Human-in-the-Loop Paradigm with Memory-Augmented AdaptationYuqing Huang, Guotian Zeng, Zhenqiao Yuan, Zhenyu He et al.CVPR 2026
Builds on28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
Related papers
- TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long VideoJinyuan Qu, Hongyang Li, Shilong Liu, Tianhe Ren et al.ICLR 2026 · 9 citations
- MixFormerV2: Efficient Fully Transformer TrackingYutao Cui, Tianhui Song, Gangshan Wu, Limin WangNeurIPS 2023 · 193 citations
- Foreground-Aware Token Routing Vision Transformer for Real-Time Satellite Video TrackingJiahao Wang, Fang Liu, Licheng Jiao, Shuo Li et al.ICML 2026
- FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with TransformersMinguk Kang, Suha KwakCVPR 2026 · 1 citation
- Efficient Video Object Segmentation and Tracking with Recurrent Dynamic SubmodelWeidong Tang, Zhiyuan Liang, Xinyan Wan, Chen Zhu et al.CVPR 2026 · 2 citations
