Exploring Historical Information for RGBE Visual Tracking with Mamba
Chuanyu Sun, Jiqing Zhang, Yang Wang, Huilin Ge, Qianchen Xia, Baocai Yin, Xin Yang
Abstract
Combining the advantages of conventional and event cameras for robust visual tracing has drawn extensive interest. However, existing tracking approaches heavily engage in complex cross-modal fusion modules, leading to higher computational complexity and training challenges. Besides, these methods generally ignore the effective integration of historical information, which is crucial to grasping the change in the target's appearance and motion trends. Given the recent advancements in Mamba's long-range modeling and linear complexity, we explore its potential in addressing the above issues in RGBE tracking tasks. Specifically, we first propose an efficient fusion module based on Mamba, which utilizes a simple gate-based interaction scheme to achieve effective modality-selective fusion. This module can be seamlessly integrated into the encoding layer of prevalent Transformer-based backbones. Moreover, we further present a novel historical decoder that leverages Mamba's advanced long sequence modeling to effectively capture the target appearance changes with autoregressive queries. Extensive experiments show that our proposed approach achieves state-of-the-art performance on multiple challenging short-term and long-term RGBE benchmarks. Besides, the effectiveness of each key Mamba-based component of our approach is evidenced by our thorough ablation study.Code will be released at: https://github . com/scy0712/MamTrack
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2040b950-0670-45ff-b558-80a42e2a22c6Cited by top-tier papers3
- SEATrack: Simple, Efficient, and Adaptive Multimodal TrackerJunbin Su, Ziteng Xue, Shihui Zhang, Kun Chen et al.CVPR 2026 · 3 citations
- Tracking through Severe Occlusion via Event-Derived Transient CuesHao Dong, Yujin Liu, Haoyue Liu, Zhenyu Wang et al.CVPR 2026
- Seeing the Unseen: Zooming in the Dark with Event CamerasDachun Kai, Zeyu Xiao, Huyue Zhu, Jiaxiao Wang et al.AAAI 2026
Builds on27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
Related papers
- Exploiting All Mamba Fusion for Efficient RGB-D TrackingGe Ying, Dawei Zhang, Chengzhuan Yang, Wei Liu et al.AAAI 2026
- RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion MambaAndong Lu, Wanyu Wang, Chenglong Li, Jin Tang et al.AAAI 2025 · 22 citations
- MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language TrackingXinqi Liu, Li Zhou, Zikun Zhou, Jianqiu Chen et al.CVPR 2025
- Exploiting Multimodal Spatial-temporal Patterns for Video Object TrackingXiantao Hu, Ying Tai, Xu Zhao, Chen Zhao et al.AAAI 2025 · 65 citations
- MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic SegmentationFuqiang Gu, Yuanke Li, Xianlei Long, Kangping Ji et al.AAAI 2026 · 1 citation
