Event Stream-Based Visual Object Tracking: A High-Resolution Benchmark Dataset and A Novel Baseline
Xiao Wang, Shiao Wang, Chuanming Tang, Lin Zhu, Bo Jiang, Yonghong Tian, Jin Tang
摘要
Tracking with bio-inspired event cameras has garnered increasing interest in recent years. Existing works either utilize aligned RGB and event data for accurate tracking or directly learn an event-based tracker. The former incurs higher inference costs while the latter may be susceptible to the impact of noisy events or sparse spatial resolution. In this paper, we propose a novel hierarchical knowledge distillation framework that can fully utilize multimodal / multi-view information during training to facilitate knowledge transfer, enabling us to achieve high-speed and low-latency visual tracking during testing by using only event signals. Specifically, a teacher Transformer-based multimodal tracking framework is first trained by feeding the RGB frame and event stream simultaneously. Then, we design a new hierarchical knowledge distillation strategy which includes pairwise similarity, feature representation, and response maps-based knowledge distillation to guide the learning of the student Transformer network. In particular, since existing event-based tracking datasets are all low-resolution (346 × 260), we propose the first large-scale high-resolution (1280 × 720) dataset named EventVOT. It contains 1141 videos and covers a wide range of categories such as pedestrians, vehicles, UAVs, ping pong, etc. Ex-tensive experiments on both low-resolution (FE240hz, Vi-sEvent, COESOT), and our newly proposed high-resolution EventVOT dataset fully validated the effectiveness of our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Global Structure-Aware Diffusion Process for Low-light Image EnhancementJinhui Hou, Zhiyu Zhu, Junhui Hou, Hui Liu 等NeurIPS 2023 · 被引用 280 次
- E-Motion: Future Motion Simulation via Event Sequence DiffusionSong Wu, Zhiyu Zhu, Junhui Hou, Guangming Shi 等NeurIPS 2024 · 被引用 14 次
- Revisiting motion information for RGB-Event tracking with MOT philosophyTianlu Zhang, Kurt Debattista, Qiang Zhang, Guiguang Ding 等NeurIPS 2024 · 被引用 11 次
- Event-Based Tiny Object Detection: A Benchmark Dataset and BaselineNuo Chen, Chao Xiao, Yimian Dai, Shiman He 等ICCV 2025 · 被引用 10 次
- XTrack: Multimodal Training Boosts RGB-X Video Object TrackersYuedong Tan, Zongwei Wu, Yuqian Fu, Zhuyun Zhou 等ICCV 2025 · 被引用 10 次
它引用的顶会 Paper16
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 被引用 746 次
- Transforming Model Prediction for TrackingChristoph Mayer, Martin Danelljan, Goutam Bhat, Matthieu Paul 等CVPR 2022 · 被引用 399 次
- Learn to Match: Automatic Matching Network Design for Visual TrackingZhipeng Zhang, Yihao Liu, Xiao Wang, Bing Li 等ICCV 2021 · 被引用 224 次
相关 Paper
- EvDistill: Asynchronous Events To End-Task Learning via Bidirectional Reconstruction-Guided Cross-Modal Knowledge DistillationLin Wang, Yujeong Chae, Sung-Hoon Yoon, Tae-Kyun Kim 等CVPR 2021
- Enhanced Event-Based Dense Stereo via Cross-Sensor Knowledge DistillationHaihao Zhang, Yunjian Zhang, Jianing Li, Lin Zhu 等ICCV 2025 · 被引用 1 次
- Frame-Event Alignment and Fusion Network for High Frame Rate TrackingJiqing Zhang, Yuanchen Wang, Wenxi Liu, Meng Li 等CVPR 2023
- Learning to Detect Objects with a 1 Megapixel Event CameraEtienne Perot, Pierre de Tournemire, Davide Nitti, Jonathan Masci 等NeurIPS 2020 · 被引用 381 次
- Recurrent Vision Transformers for Object Detection with Event CamerasMathias Gehrig, Davide ScaramuzzaCVPR 2023
