Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking
Shilei Wang, Pujian Lai, Dong Gao, Jifeng Ning, Gong Cheng
摘要
Most existing multi-modal trackers adopt uniform fusion strategies, overlooking the inherent differences between modalities. Moreover, they propagate temporal information through mixed tokens, leading to entangled and less discriminative temporal representations. To address these limitations, we propose MDTrack, a novel framework for modalityaware fusion and decoupled temporal propagation in multimodal object tracking. Specifically, for modality-aware fusion, we allocate dedicated experts to each modality (Infrared, Event, Depth, and RGB) to process their respective representations. The gating mechanism within the Mixture of Experts (MoE) then dynamically selects the optimal experts based on the input features, enabling adaptive and modality-specific fusion. For decoupled temporal propagation, we introduce two separate State Space Model (SSM) structures to independently store and update the hidden states h of the RGB and X-modal streams, effectively capturing their distinct temporal information. To ensure synergy between the two temporal representations, we incorporate a set of cross-attentions between the input features of the two SSMs, facilitating implicit information exchange. The resulting temporally enriched features are then integrated into the backbone via another set of cross-attention, enhancing MDTrack's ability to leverage temporal information. Extensive experiments demonstrate the effectiveness of our proposed method. Both MDTrack-S (Modality-Specific Training) and MDTrack-U (Unified-Modality Training) achieve state-of-the-art performance across five multimodal tracking benchmarks. The code is publicly available at https://github.com/wsumel/MDTrack
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of ExpertsBasil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton 等NeurIPS 2022 · 被引用 359 次
- ODTrack: Online Dense Temporal Token Learning for Visual TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo 等AAAI 2024 · 被引用 247 次
- Prompting for Multi-Modal TrackingJinyu Yang, Zhe Li, Feng Zheng, Ales Leonardis 等ACM MM 2022 · 被引用 167 次
相关 Paper
- UETrack: A Unified and Efficient Framework for Single Object TrackingBen Kang, Jie Zhao, Xin Chen, Wanting Geng 等CVPR 2026 · 被引用 2 次
- CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingHao Li, Yuhao Wang, Xiantao Hu, Wenning Hao 等AAAI 2026 · 被引用 4 次
- Single-Model and Any-Modality for Video Object TrackingZongwei Wu, Jilai Zheng, Xiangxuan Ren, Florin-Alexandru Vasluianu 等CVPR 2024 · 被引用 78 次
- Unified Multimodal Visual Tracking with Dual Mixture-of-ExpertsLingyi Hong, Jinglun Li, Xinyu Zhou, Kaixun Jiang 等ICML 2026
- Exploiting All Mamba Fusion for Efficient RGB-D TrackingGe Ying, Dawei Zhang, Chengzhuan Yang, Wei Liu 等AAAI 2026
