Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking
You Wu, Xucheng Wang, Xiangyang Yang, Mengyuan Liu, Dan Zeng, Hengzhou Ye, Shuiwang Li
Abstract
Single-stream architectures using Vision Transformer (ViT) backbones show great potential for real-time UAV tracking recently. However, frequent occlusions from obstacles like buildings and trees expose a major drawback: these models often lack strategies to handle occlusions effectively. New methods are needed to enhance the occlusion resilience of single-stream ViT models in aerial tracking. In this work, we propose to learn Occlusion-Robust Representations (ORR) based on ViTs for UAV tracking by enforcing an invariance of the feature representation of a target with respect to random masking operations modeled by a spatial Cox process. Hopefully, this random masking approximately simulates target occlusions, thereby enabling us to learn ViTs that are robust to target occlusion for UAV tracking. This framework is termed ORTrack. Additionally, to facilitate real-time applications, we propose an Adaptive Feature-Based Knowledge Distillation (AFKD) method to create a more compact tracker, which adaptively mimics the behavior of the teacher model ORTrack according to the task's difficulty. This student model, dubbed ORTrack-D, retains much of ORTrack's performance while offering higher efficiency. Extensive experiments on multiple benchmarks validate the effectiveness of our method, demonstrating its state-of-the-art performance. Codes is available at https://github.com/wuyou3474/ORTrack .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cb0e6d2-a97c-4be0-9ebf-b8f6be46a0a6Cited by top-tier papers9
- Instance-level Visual Active Tracking with Occlusion-Aware PlanningHaowei Sun, Kai Zhou, Hao Gao, Shiteng Zhang et al.CVPR 2026 · 4 citations
- Rethinking Occlusion Modeling for UAV TrackingJian Zhang, Xincheng Yu, Yi LinCVPR 2026
- Exploring Reliable Spatiotemporal Dependencies for Efficient Visual TrackingJunze Shi, Yang Yu, Jian Shi, Haibo LuoAAAI 2026
- Tracking through Severe Occlusion via Event-Derived Transient CuesHao Dong, Yujin Liu, Haoyue Liu, Zhenyu Wang et al.CVPR 2026
- UAST: Unified Active Search and Tracking for Arbitrary Targets with UAVsLiang Qin, Min Wang, Xingyu Lu, Aowen Qiu et al.CVPR 2026
Builds on34
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
Related papers
- Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV TrackingYongxin Li, Mengyuan Liu, You Wu, Xucheng Wang et al.ICML 2024 · 63 citations
- Dual-branch Distilled Transformer for Efficient Asymmetric UAV TrackingHongtao Yang, Bineng Zhong, Qihua Liang, Yaozong Zheng et al.CVPR 2026
- Similarity-Guided Layer-Adaptive Vision Transformer for UAV TrackingChaocan Xue, Bineng Zhong, Qihua Liang, Yaozong Zheng et al.CVPR 2025
- Adaptive and Background-Aware Vision Transformer for Real-Time UAV TrackingShuiwang Li, Xiangxyang Yang, Dan Zeng, Xucheng WangICCV 2023 · 74 citations
- Foreground-Aware Token Routing Vision Transformer for Real-Time Satellite Video TrackingJiahao Wang, Fang Liu, Licheng Jiao, Shuo Li et al.ICML 2026
