XTrack: Multimodal Training Boosts RGB-X Video Object Trackers
Yuedong Tan, Zongwei Wu, Yuqian Fu, Zhuyun Zhou, Guolei Sun, Eduard Zamfir, Chao Ma, Danda Pani Paudel, Luc Van Gool, Radu Timofte
Abstract
Multimodal sensing has proven valuable for visual tracking, as different sensor types offer unique strengths in handling one specific challenging scene where object appearance varies. While a generalist model capable of leveraging all modalities would be ideal, development is hindered by data sparsity, typically in practice, only one modality is available at a time. Therefore, it is crucial to ensure and achieve that knowledge gained from multimodal sensing - such as identifying relevant features and regions is effectively shared, even when certain modalities are unavailable at inference. We venture with a simple assumption: similar samples across different modalities have more knowledge to share than otherwise. To implement this, we employ a classifier with weak loss tasked with distinguishing between modalities. More specifically, if the classifier “fails” to accurately identify the modality of the given sample, this signals an opportunity for cross-modal knowledge sharing. Intuitively, knowledge transfer is facilitated whenever a sample from one modality is sufficiently close and aligned with another. Technically, we achieve this by routing samples from one modality to the expert of the others, within a mixture-of-experts framework designed for multimodal video object tracking. During the inference, the expert of the respective modality is chosen, which we show to benefit from the multimodal knowledge available during training, thanks to the proposed method. Through the exhaustive experiments that use only paired RGB-E, RGBD, and RGB-T during training, we showcase the benefit of the proposed method for RGB-X tracker during inference, with an average precision improvement over the current SOTA. The source code is publicly available at https://github.com/supertyd/XTrack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f388bb12-3f3e-453d-b0e6-faf6206f7f25Cited by top-tier papers14
- CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud TrackingSifan Zhou, Yichao Cao, Jiahao Nie, Yuqian Fu et al.AAAI 2026 · 9 citations
- INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction TuningWujian Peng, Lingchen Meng, Yitong Chen, Yiweng Xie et al.NeurIPS 2025 · 7 citations
- ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric PerspectivesYuqian Fu, Runze Wang, Bin Ren, Guolei Sun et al.ICCV 2025 · 5 citations
- CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingHao Li, Yuhao Wang, Xiantao Hu, Wenning Hao et al.AAAI 2026 · 4 citations
- SEATrack: Simple, Efficient, and Adaptive Multimodal TrackerJunbin Su, Ziteng Xue, Shihui Zhang, Kun Chen et al.CVPR 2026 · 3 citations
Builds on39
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric TasksWenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu et al.NeurIPS 2023 · 725 citations
Related papers
- Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object TrackingShilei Wang, Pujian Lai, Dong Gao, Jifeng Ning et al.AAAI 2026
- Unified Multimodal Visual Tracking with Dual Mixture-of-ExpertsLingyi Hong, Jinglun Li, Xinyu Zhou, Kaixun Jiang et al.ICML 2026
- What You Have is What You Track: Adaptive and Robust Multimodal TrackingYuedong Tan, Jiawei Shao, Eduard Zamfir, Ruanjun Li et al.ICCV 2025 · 5 citations
- Prompting for Multi-Modal TrackingJinyu Yang, Zhe Li, Feng Zheng, Ales Leonardis et al.ACM MM 2022 · 167 citations
- Bi-directional Adapter for Multimodal TrackingBing Cao, Junliang Guo, Pengfei Zhu, Qinghua HuAAAI 2024 · 153 citations
