LoRATv2: Enabling Low-Cost Temporal Modeling in One-Stream Trackers
Liting Lin, Heng Fan, Zhipeng Zhang, Yuqing Huang, Yaowei Wang, Yong Xu, Haibin Ling
摘要
Transformer-based algorithms, such as LoRAT, have significantly enhanced objecttracking performance. However, these approaches rely on a standard attention mechanism, which incurs quadratic token complexity, making real-time inference computationally expensive. In this paper, we introduce LoRATv2, a novel tracking framework that addresses these limitations with three main contributions. First, LoRATv2 integrates frame-wise causal attention, which ensures full selfattention within each frame while enabling causal dependencies across frames, significantly reducing computational overhead. Moreover, key-value (KV) caching is employed to efficiently reuse past embeddings for further speedup. Second, building on LoRAT's parameter-efficient fine-tuning, we propose Stream-Specific LoRA Adapters (SSLA). As frame-wise causal attention introduces asymmetry in how streams access temporal information, SSLA assigns dedicated LoRA modules to the template and each search stream, with the main ViT backbone remaining frozen. This allows specialized adaptation for each stream's role in temporal tracking. Third, we introduce a two-phase progressive training strategy, which first trains a single-search-frame tracker and then gradually extends it to multi-searchframe inputs by introducing additional LoRA modules. This curriculum-based learning paradigm improves long-term tracking while maintaining training efficiency. In extensive experiments on multiple benchmarks, LoRATv2 achieves state-of-the-art performance, substantially improved efficiency, and a superior performance-to-FLOPs ratio over state-of-the-art trackers. The code is available at https://github.com/LitingLin/LoRATv2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Drift-Resilient Temporal Priors for Visual TrackingYuqing Huang, Liting Lin, Weijun Zhuang, Zhenyu He 等CVPR 2026 · 被引用 1 次
- Interactive Tracking: A Human-in-the-Loop Paradigm with Memory-Augmented AdaptationYuqing Huang, Guotian Zeng, Zhenqiao Yuan, Zhenyu He 等CVPR 2026
它引用的顶会 Paper28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
相关 Paper
- TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long VideoJinyuan Qu, Hongyang Li, Shilong Liu, Tianhe Ren 等ICLR 2026 · 被引用 9 次
- MixFormerV2: Efficient Fully Transformer TrackingYutao Cui, Tianhui Song, Gangshan Wu, Limin WangNeurIPS 2023 · 被引用 193 次
- Foreground-Aware Token Routing Vision Transformer for Real-Time Satellite Video TrackingJiahao Wang, Fang Liu, Licheng Jiao, Shuo Li 等ICML 2026
- FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with TransformersMinguk Kang, Suha KwakCVPR 2026 · 被引用 1 次
- Efficient Video Object Segmentation and Tracking with Recurrent Dynamic SubmodelWeidong Tang, Zhiyuan Liang, Xinyan Wan, Chen Zhu 等CVPR 2026 · 被引用 2 次
