Foreground-Aware Token Routing Vision Transformer for Real-Time Satellite Video Tracking
Jiahao Wang, Fang Liu, Licheng Jiao, Shuo Li, Hao Wang, Lingling Li, Xinyi Wang, Xu Liu
摘要
Real-time satellite video tracking poses distinct challenges, including accommodating high spatial-temporal resolution, dynamic backgrounds, and constrained onboard computational resources. While Discriminative Correlation Filter (DCF)-based methods offer high-speed inference, they suffer from limited accuracy. In contrast, Vision Transformer (ViT)-based trackers achieve strong performance by unifying representation and aggregation in a single-stream design, yet their heavy computational footprint limits practical deployment in real-time satellite scenarios. In this work, we present FATrack, a novel tracking framework that effectively balances tracking accuracy and computational efficiency. At its core is FA-ViT, a lightweight Vision Transformer backbone that introduces foreground-aware token routing, enabling the model to concentrate computation on target-relevant regions while suppressing redundancy. To mitigate semantic degradation caused by token sparsification, we propose the Adaptive Scatter Module (ASM), which selectively reinforces informative tokens via joint spatial-channel attention and sparse structural propagation, thereby enhancing both semantic fidelity and spatial coherence. By synergistically integrating FA-ViT and ASM, FATrack forms a unified architecture that delivers real-time performance with significantly improved tracking precision. Extensive evaluations on multiple satellite video benchmarks demonstrate that FATrack surpasses existing real-time trackers in accuracy and achieves inference efficiency comparable to DCF-based methods, highlighting its potential for practical deployment in large-scale aerial video tracking systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 被引用 2,162 次
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
相关 Paper
- Adaptive and Background-Aware Vision Transformer for Real-Time UAV TrackingShuiwang Li, Xiangxyang Yang, Dan Zeng, Xucheng WangICCV 2023 · 被引用 74 次
- Similarity-Guided Layer-Adaptive Vision Transformer for UAV TrackingChaocan Xue, Bineng Zhong, Qihua Liang, Yaozong Zheng 等CVPR 2025
- Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV TrackingYongxin Li, Mengyuan Liu, You Wu, Xucheng Wang 等ICML 2024 · 被引用 63 次
- Learning Occlusion-Robust Vision Transformers for Real-Time UAV TrackingYou Wu, Xucheng Wang, Xiangyang Yang, Mengyuan Liu 等CVPR 2025
- Compact Transformer Tracker with Correlative Masked ModelingZikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen 等AAAI 2023 · 被引用 136 次
