Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking
You Wu, Xucheng Wang, Xiangyang Yang, Mengyuan Liu, Dan Zeng, Hengzhou Ye, Shuiwang Li
摘要
Single-stream architectures using Vision Transformer (ViT) backbones show great potential for real-time UAV tracking recently. However, frequent occlusions from obstacles like buildings and trees expose a major drawback: these models often lack strategies to handle occlusions effectively. New methods are needed to enhance the occlusion resilience of single-stream ViT models in aerial tracking. In this work, we propose to learn Occlusion-Robust Representations (ORR) based on ViTs for UAV tracking by enforcing an invariance of the feature representation of a target with respect to random masking operations modeled by a spatial Cox process. Hopefully, this random masking approximately simulates target occlusions, thereby enabling us to learn ViTs that are robust to target occlusion for UAV tracking. This framework is termed ORTrack. Additionally, to facilitate real-time applications, we propose an Adaptive Feature-Based Knowledge Distillation (AFKD) method to create a more compact tracker, which adaptively mimics the behavior of the teacher model ORTrack according to the task's difficulty. This student model, dubbed ORTrack-D, retains much of ORTrack's performance while offering higher efficiency. Extensive experiments on multiple benchmarks validate the effectiveness of our method, demonstrating its state-of-the-art performance. Codes is available at https://github.com/wuyou3474/ORTrack .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Instance-level Visual Active Tracking with Occlusion-Aware PlanningHaowei Sun, Kai Zhou, Hao Gao, Shiteng Zhang 等CVPR 2026 · 被引用 4 次
- Rethinking Occlusion Modeling for UAV TrackingJian Zhang, Xincheng Yu, Yi LinCVPR 2026
- Exploring Reliable Spatiotemporal Dependencies for Efficient Visual TrackingJunze Shi, Yang Yu, Jian Shi, Haibo LuoAAAI 2026
- Tracking through Severe Occlusion via Event-Derived Transient CuesHao Dong, Yujin Liu, Haoyue Liu, Zhenyu Wang 等CVPR 2026
- UAST: Unified Active Search and Tracking for Arbitrary Targets with UAVsLiang Qin, Min Wang, Xingyu Lu, Aowen Qiu 等CVPR 2026
它引用的顶会 Paper34
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
相关 Paper
- Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV TrackingYongxin Li, Mengyuan Liu, You Wu, Xucheng Wang 等ICML 2024 · 被引用 63 次
- Dual-branch Distilled Transformer for Efficient Asymmetric UAV TrackingHongtao Yang, Bineng Zhong, Qihua Liang, Yaozong Zheng 等CVPR 2026
- Similarity-Guided Layer-Adaptive Vision Transformer for UAV TrackingChaocan Xue, Bineng Zhong, Qihua Liang, Yaozong Zheng 等CVPR 2025
- Adaptive and Background-Aware Vision Transformer for Real-Time UAV TrackingShuiwang Li, Xiangxyang Yang, Dan Zeng, Xucheng WangICCV 2023 · 被引用 74 次
- Foreground-Aware Token Routing Vision Transformer for Real-Time Satellite Video TrackingJiahao Wang, Fang Liu, Licheng Jiao, Shuo Li 等ICML 2026
