RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba
Andong Lu, Wanyu Wang, Chenglong Li, Jin Tang, Bin Luo
摘要
Existing RGBT tracking methods often design various interaction models to perform cross-modal fusion of each layer, but can not execute the feature interactions among all layers, which plays a critical role in robust multimodal representation, due to large computational burden. To address this issue, this paper presents a novel All-layer multimodal Interaction Network, named AINet, which performs efficient and effective feature interactions of all modalities and layers in a progressive fusion Mamba, for robust RGBT tracking. Even though modality features in different layers are known to contain different cues, it is always challenging to build multimodal interactions in each layer due to struggling in balancing interaction capabilities and efficiency. Meanwhile, considering that the feature discrepancy between RGB and thermal modalities reflects their complementary information to some extent, we design a Difference-based Fusion Mamba (DFM) to achieve enhanced fusion of different modalities with linear complexity. When interacting with features from all layers, a huge number of token sequences (3840 tokens in this work) are involved and the computational burden is thus large. To handle this problem, we design an Order-dynamic Fusion Mamba (OFM) to execute efficient and effective feature interactions of all layers by dynamically adjusting the scan order of different layers in Mamba. Extensive experiments on four public RGBT tracking datasets show that AINet achieves leading performance against existing state-of-the-art methods. We will release the code upon acceptance of the paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Adaptive Depth Lightweight RGB-T Tracking with Holistic Token RoutingTian Ding, Hongtao Yang, Liangtao Shi, Jun Li 等CVPR 2026 · 被引用 3 次
- RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented GenerationHao Li, Yuhao Wang, Wenning Hao, Pingping Zhang 等CVPR 2026 · 被引用 2 次
- Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT TrackingAndong Lu, Ziyi Zha, Jiandong Jin, Shihao Li 等CVPR 2026 · 被引用 2 次
- MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video ParsingLangyu Wang, Bingke Zhu, Yingying Chen, Yiyuan Zhang 等ICCV 2025 · 被引用 1 次
- Tracking and Segmenting Anything in Any ModalityTianlu Zhang, Qiang Zhang, Guiguang Ding, Jungong HanAAAI 2026
它引用的顶会 Paper25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
相关 Paper
- CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingHao Li, Yuhao Wang, Xiantao Hu, Wenning Hao 等AAAI 2026 · 被引用 4 次
- Simplifying Cross-modal Interaction via Modality-Shared Features for RGBT TrackingLiqiu Chen, Yuqing Huang, Hengyu Li, Zikun Zhou 等ACM MM 2024 · 被引用 2 次
- Attribute-Based Progressive Fusion Network for RGBT TrackingYun Xiao, Mengmeng Yang, Chenglong Li, Lei Liu 等AAAI 2022 · 被引用 218 次
- All-Day Multi-Camera Multi-Target TrackingHuijie Fan, Yu Qiao, Yihao Zhen, Tinghui Zhao 等CVPR 2025
- Exploring Historical Information for RGBE Visual Tracking with MambaChuanyu Sun, Jiqing Zhang, Yang Wang, Huilin Ge 等CVPR 2025
