Progressive Multi-cue Alignment for Unaligned RGBT Tracking
Jiandong Jin, Chenglong Li, Hao Feng, Andong Lu, Lili Huang, Jin Tang
Abstract
Unaligned RGBT tracking aims to achieve robust target localization across spatially misaligned RGB and thermal infrared (TIR) videos, a crucial challenge for deploying RGBT tracking in real-world scenarios. Existing methods often calculate all cross-modal alignment parameters (i.e., spatial shift and scale change) simultaneously, but suffer from two major limitations. 1) They are difficult to adapt to different degrees of unaligned difficulty during tracking. 2) They usually require complex models to handle challenging scenarios, resulting in a large computational burden. To overcome these limitations, we propose a novel Progressive Multi-cue Alignment framework called PMATrack, which disentangles the calculation of cross-modal alignment parameters in a progressive manner and dynamically selects appropriate cues to handle different challenges, thereby enabling robust and efficient unaligned RGBT tracking. In particular, PMATrack divides the cross-modal alignment parameter estimation into three stages to progressively perform center offset computation, scale transformation estimation, and global refinement. At each stage, we design a difficulty-aware router to adaptively select the appropriate alignment expert based on the cross-modal alignment complexity, thereby reducing computational redundancy. In addition, we build a high-quality video benchmark called MUART244 to facilitate the comprehensive evaluation of different unaligned RGBT tracking algorithms. Extensive experiments demonstrate the outstanding performance of PMATrack against existing state-of-the-art methods. The code and dataset will be available at https://github. com/NOP1224/Unaligned_RGBT_Tracking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0b4296b-0365-4d4e-8aba-38327bebc5f4Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Transforming Model Prediction for TrackingChristoph Mayer, Martin Danelljan, Goutam Bhat, Matthieu Paul et al.CVPR 2022 · 399 citations
- Visible-Thermal UAV Tracking: A Large-Scale Benchmark and New BaselinePengyu Zhang, Jie Zhao, Dong Wang, Huchuan Lu et al.CVPR 2022 · 225 citations
- Attribute-Based Progressive Fusion Network for RGBT TrackingYun Xiao, Mengmeng Yang, Chenglong Li, Lei Liu et al.AAAI 2022 · 218 citations
Related papers
- Unaligned UAV RGBT Tracking: A Largescale Benchmark and a Novel ApproachYun Xiao, Yuhang Wang, Jiandong Jin, Wankang Zhang et al.AAAI 2026
- CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingHao Li, Yuhao Wang, Xiantao Hu, Wenning Hao et al.AAAI 2026 · 4 citations
- AlignTrack: Top-Down Spatiotemporal Resolution Alignment for RGB-Event Visual TrackingChuanyu Sun, Jiqing Zhang, Yang Wang, Yuanchen Wang et al.AAAI 2026
- Quality-Aware RGBT Tracking via Supervised Reliability Learning and Weighted Residual GuidanceLei Liu, Chenglong Li, Yun Xiao, Jin TangACM MM 2023 · 36 citations
- RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented GenerationHao Li, Yuhao Wang, Wenning Hao, Pingping Zhang et al.CVPR 2026 · 2 citations
