Cross-Modal Object Tracking: Modality-Aware Representations and a Unified Benchmark
Chenglong Li, Tianhao Zhu, Lei Liu, Xiaonan Si, Zilin Fan, Sulan Zhai
Abstract
In many visual systems, visual tracking often bases on RGB image sequences, in which some targets are invalid in lowlight conditions, and tracking performance is thus affected significantly. Introducing other modalities such as depth and infrared data is an effective way to handle imaging limitations of individual sources, but multi-modal imaging platforms usually require elaborate designs and cannot be applied in many real-world applications at present. Near-infrared (NIR) imaging becomes an essential part of many surveillance cameras, whose imaging is switchable between RGB and NIR based on the light intensity. These two modalities are heterogeneous with very different visual properties and thus bring big challenges for visual tracking. However, existing works have not studied this challenging problem. In this work, we address the cross-modal object tracking problem and contribute a new video dataset, including 654 cross-modal image sequences with over 481K frames in total, and the average video length is more than 735 frames. To promote the research and development of cross-modal object tracking, we propose a new algorithm, which learns the modality-aware target representation to mitigate the appearance gap between RGB and NIR modalities in the tracking process. It is plugand-play and could thus be flexibly embedded into different tracking frameworks. Extensive experiments on the dataset are conducted, and we demonstrate the effectiveness of the proposed algorithm in two representative tracking frameworks against 17 state-of-the-art tracking methods. We will release the dataset for free academic usage, dataset download link and code will be released soon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext add83adb-e29d-437a-ac45-36c2061ef70bBuilds on5
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- GlobalTrack: A Simple and Strong Baseline for Long-Term TrackingLianghua Huang, Xin Zhao, Kaiqi HuangAAAI 2020 · 278 citations
- GradNet: Gradient-Guided Network for Visual Object TrackingPeixia Li, Boyu Chen, Wanli Ouyang, Dong Wang et al.ICCV 2019 · 255 citations
- 'Skimming-Perusal' Tracking: A Framework for Real-Time and Robust Long-Term TrackingBin Yan, Haojie Zhao, Dong Wang, Huchuan Lu et al.ICCV 2019 · 177 citations
- High-Performance Long-Term Tracking With Meta-UpdaterKenan Dai, Yunhua Zhang, Dong Wang, Jianhua Li et al.CVPR 2020
Related papers
- All-Day Multi-Camera Multi-Target TrackingHuijie Fan, Yu Qiao, Yihao Zhen, Tinghui Zhao et al.CVPR 2025
- Unaligned UAV RGBT Tracking: A Largescale Benchmark and a Novel ApproachYun Xiao, Yuhang Wang, Jiandong Jin, Wankang Zhang et al.AAAI 2026
- Multi-Task Driven Feature Models for Thermal Infrared TrackingQiao Liu, Xin Li, Zhenyu He, Nana Fan et al.AAAI 2020 · 73 citations
- Low-light Invariant Representation Learning for Visible-Infrared Person Re-identificationDengwen Wang, Guanyu Xing, Yanli LiuACM MM 2025 · 3 citations
- Simplifying Cross-modal Interaction via Modality-Shared Features for RGBT TrackingLiqiu Chen, Yuqing Huang, Hengyu Li, Zikun Zhou et al.ACM MM 2024 · 2 citations
