Revisiting motion information for RGB-Event tracking with MOT philosophy
Tianlu Zhang, Kurt Debattista, Qiang Zhang, Guiguang Ding, Jungong Han
Abstract
RGB-Event single object tracking (SOT) aims to leverage the merits of RGB and event data to achieve higher performance. However, existing frameworks focus on exploring complementary appearance information within multi-modal data, and struggle to address the association problem of targets and distractors in the temporal domain using motion information from the event stream. In this paper, we introduce the Multi-Object Tracking (MOT) philosophy into RGB-E SOT to keep track of targets as well as distractors by using both RGB and event data, thereby improving the robustness of the tracker. Specifically, an appearance model is employed to predict the initial candidates. Subsequently, the initially predicted tracking results, in combination with the RGB-E features, are encoded into appearance and motion embeddings, respectively. Furthermore, a Spatial-Temporal Transformer Encoder is proposed to model the spatial-temporal relationships and learn discriminative features for each candidate through guidance of the appearance-motion embeddings. Simultaneously, a Dual-Branch Transformer Decoder is designed to adopt such motion and appearance information for candidate matching, thus distinguishing between targets and distractors. The proposed method is evaluated on multiple benchmark datasets and achieves state-of-the-art performance on all the datasets tested.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fa48542-68e6-404e-a986-c307fd4cd016Cited by top-tier papers2
- Tracking through Severe Occlusion via Event-Derived Transient CuesHao Dong, Yujin Liu, Haoyue Liu, Zhenyu Wang et al.CVPR 2026
- AlignTrack: Top-Down Spatiotemporal Resolution Alignment for RGB-Event Visual TrackingChuanyu Sun, Jiqing Zhang, Yang Wang, Yuanchen Wang et al.AAAI 2026
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- Transforming Model Prediction for TrackingChristoph Mayer, Martin Danelljan, Goutam Bhat, Matthieu Paul et al.CVPR 2022 · 399 citations
- Learning Target Candidate Association to Keep Track of What Not to TrackChristoph Mayer, Martin Danelljan, Danda Pani Paudel, Luc Van GoolICCV 2021 · 356 citations
Related papers
- Online Multiple Object Tracking With Cross-Task SynergySong Guo, Jingya Wang, Xinchao Wang, Dacheng TaoCVPR 2021
- MeMOT: Multi-Object Tracking with MemoryJiarui Cai, Mingze Xu, Wei Li, Yuanjun Xiong et al.CVPR 2022 · 216 citations
- MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object TrackingRuopeng Gao, Limin WangICCV 2023 · 143 citations
- Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object TrackingShilei Wang, Pujian Lai, Dong Gao, Jifeng Ning et al.AAAI 2026
- Dual-Path Temporal Decoder for End-to-End Multi-Object TrackingHyunseop Kim, Juheon Jeong, Hanul Kim, Yeong Jun KohNeurIPS 2025 · 4 citations
