Event Stream-Based Visual Object Tracking: A High-Resolution Benchmark Dataset and A Novel Baseline
Xiao Wang, Shiao Wang, Chuanming Tang, Lin Zhu, Bo Jiang, Yonghong Tian, Jin Tang
Abstract
Tracking with bio-inspired event cameras has garnered increasing interest in recent years. Existing works either utilize aligned RGB and event data for accurate tracking or directly learn an event-based tracker. The former incurs higher inference costs while the latter may be susceptible to the impact of noisy events or sparse spatial resolution. In this paper, we propose a novel hierarchical knowledge distillation framework that can fully utilize multimodal / multi-view information during training to facilitate knowledge transfer, enabling us to achieve high-speed and low-latency visual tracking during testing by using only event signals. Specifically, a teacher Transformer-based multimodal tracking framework is first trained by feeding the RGB frame and event stream simultaneously. Then, we design a new hierarchical knowledge distillation strategy which includes pairwise similarity, feature representation, and response maps-based knowledge distillation to guide the learning of the student Transformer network. In particular, since existing event-based tracking datasets are all low-resolution (346 × 260), we propose the first large-scale high-resolution (1280 × 720) dataset named EventVOT. It contains 1141 videos and covers a wide range of categories such as pedestrians, vehicles, UAVs, ping pong, etc. Ex-tensive experiments on both low-resolution (FE240hz, Vi-sEvent, COESOT), and our newly proposed high-resolution EventVOT dataset fully validated the effectiveness of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9797fd9-7e53-4fed-8dcd-f8182aa01290Cited by top-tier papers22
- Global Structure-Aware Diffusion Process for Low-light Image EnhancementJinhui Hou, Zhiyu Zhu, Junhui Hou, Hui Liu et al.NeurIPS 2023 · 280 citations
- E-Motion: Future Motion Simulation via Event Sequence DiffusionSong Wu, Zhiyu Zhu, Junhui Hou, Guangming Shi et al.NeurIPS 2024 · 14 citations
- Revisiting motion information for RGB-Event tracking with MOT philosophyTianlu Zhang, Kurt Debattista, Qiang Zhang, Guiguang Ding et al.NeurIPS 2024 · 11 citations
- Event-Based Tiny Object Detection: A Benchmark Dataset and BaselineNuo Chen, Chao Xiao, Yimian Dai, Shiman He et al.ICCV 2025 · 10 citations
- XTrack: Multimodal Training Boosts RGB-X Video Object TrackersYuedong Tan, Zongwei Wu, Yuqian Fu, Zhuyun Zhou et al.ICCV 2025 · 10 citations
Builds on16
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- Transforming Model Prediction for TrackingChristoph Mayer, Martin Danelljan, Goutam Bhat, Matthieu Paul et al.CVPR 2022 · 399 citations
- Learn to Match: Automatic Matching Network Design for Visual TrackingZhipeng Zhang, Yihao Liu, Xiao Wang, Bing Li et al.ICCV 2021 · 224 citations
Related papers
- EvDistill: Asynchronous Events To End-Task Learning via Bidirectional Reconstruction-Guided Cross-Modal Knowledge DistillationLin Wang, Yujeong Chae, Sung-Hoon Yoon, Tae-Kyun Kim et al.CVPR 2021
- Enhanced Event-Based Dense Stereo via Cross-Sensor Knowledge DistillationHaihao Zhang, Yunjian Zhang, Jianing Li, Lin Zhu et al.ICCV 2025 · 1 citation
- Frame-Event Alignment and Fusion Network for High Frame Rate TrackingJiqing Zhang, Yuanchen Wang, Wenxi Liu, Meng Li et al.CVPR 2023
- Learning to Detect Objects with a 1 Megapixel Event CameraEtienne Perot, Pierre de Tournemire, Davide Nitti, Jonathan Masci et al.NeurIPS 2020 · 381 citations
- Recurrent Vision Transformers for Object Detection with Event CamerasMathias Gehrig, Davide ScaramuzzaCVPR 2023
