Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking
Yongxin Li, Mengyuan Liu, You Wu, Xucheng Wang, Xiangyang Yang, Shuiwang Li
Abstract
Single-stream architectures using Vision Transformer (ViT) backbones show great potential for real-time UAV tracking recently. However, frequent occlusions from obstacles like buildings and trees expose a major drawback: these models often lack strategies to handle occlusions effectively. New methods are needed to enhance the occlusion resilience of single-stream ViT models in aerial tracking. In this work, we propose to learn Occlusion-Robust Representations (ORR) based on ViTs for UAV tracking by enforcing an invariance of the feature representation of a target with respect to random masking operations modeled by a spatial Cox process. Hopefully, this random masking approximately simulates target occlusions, thereby enabling us to learn ViTs that are robust to target occlusion for UAV tracking. This framework is termed ORTrack. Additionally, to facilitate real-time applications, we propose an Adaptive Feature-Based Knowledge Distillation (AFKD) method to create a more compact tracker, which adaptively mimics the behavior of the teacher model ORTrack according to the task's difficulty. This student model, dubbed ORTrack-D, retains much of ORTrack's performance while offering higher efficiency. Extensive experiments on multiple benchmarks validate the effectiveness of our method, demonstrating its state-of-the-art performance. Codes is available at https://github.com/wuyou3474/ORTrack .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba0e3855-ada6-4c43-9df2-40297bfd3583Cited by top-tier papers10
- LVPTrack: High Performance Domain Adaptive UAV Tracking with Label Aligned Visual Prompt TuningHongjing Wu, Siyuan Yao, Feng Huang, Shu Wang et al.AAAI 2025 · 5 citations
- UMDATrack: Unified Multi-Domain Adaptive Tracking under Adverse Weather ConditionsSiyuan Yao, Rui Zhu, Ziqi Wang, Wenqi Ren et al.ICCV 2025 · 4 citations
- Tracking Tiny Drones Against Clutter: Large-Scale Infrared Benchmark with Motion-Centric Adaptive AlgorithmJiahao Zhang, Zongli Jiang, Jinli Zhang, Yixin Wei et al.ICCV 2025 · 3 citations
- Multi-State Tracker: Enhancing Efficient Object Tracking via Multi-State Specialization and InteractionShilei Wang, Gong Cheng, Pujian Lai, Dong Gao et al.ACM MM 2025 · 2 citations
- Dynamic Semantic-Aware Correlation Modeling for UAV TrackingXinyu Zhou, Tongxin Pan, Lingyi Hong, Pinxue Guo et al.NeurIPS 2025 · 2 citations
Builds on33
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
Related papers
- Learning Occlusion-Robust Vision Transformers for Real-Time UAV TrackingYou Wu, Xucheng Wang, Xiangyang Yang, Mengyuan Liu et al.CVPR 2025
- Rethinking Occlusion Modeling for UAV TrackingJian Zhang, Xincheng Yu, Yi LinCVPR 2026
- Dual-branch Distilled Transformer for Efficient Asymmetric UAV TrackingHongtao Yang, Bineng Zhong, Qihua Liang, Yaozong Zheng et al.CVPR 2026
- Similarity-Guided Layer-Adaptive Vision Transformer for UAV TrackingChaocan Xue, Bineng Zhong, Qihua Liang, Yaozong Zheng et al.CVPR 2025
- Adaptive and Background-Aware Vision Transformer for Real-Time UAV TrackingShuiwang Li, Xiangxyang Yang, Dan Zeng, Xucheng WangICCV 2023 · 74 citations
