Recurrent Vision Transformers for Object Detection with Event Cameras
Mathias Gehrig, Davide Scaramuzza
Abstract
We present Recurrent Vision Transformers (RVTs), a novel backbone for object detection with event cameras. Event cameras provide visual information with submillisecond latency at a high-dynamic range and with strong robustness against motion blur. These unique properties offer great potential for low-latency object detection and tracking in time-critical scenarios. Prior work in event-based vision has achieved outstanding detection performance but at the cost of substantial inference time, typically beyond 40 milliseconds. By revisiting the high-level design of recurrent vision backbones, we reduce inference time by a factor of 6 while retaining similar performance. To achieve this, we explore a multi-stage design that utilizes three key concepts in each stage: first, a convolutional prior that can be regarded as a conditional positional embedding. Second, local and dilated global self-attention for spatial feature interaction. Third, recurrent temporal feature aggregation to minimize latency while retaining temporal information. RVTs can be trained from scratch to reach state-of-the-art performance on event-based object detection -achieving an mAP of 47.2% on the Gen1 automotive dataset. At the same time, RVTs offer fast inference (< 12 ms on a T4 GPU) and favorable parameter efficiency (5× fewer than prior art). Our study brings new insights into effective design choices that can be fruitful for research beyond event-based vision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af299316-ff23-411a-9985-c74efb112346Cited by top-tier papers52
- GET: Group Event Transformer for Event-Based VisionYansong Peng, Yueyi Zhang, Zhiwei Xiong, Xiaoyan Sun et al.ICCV 2023 · 86 citations
- From Chaos Comes Order: Ordering Event Representations for Object Recognition and DetectionNikola Zubic, Daniel Gehrig, Mathias Gehrig, Davide ScaramuzzaICCV 2023 · 71 citations
- State Space Models for Event CamerasNikola Zubic, Mathias Gehrig, Davide ScaramuzzaCVPR 2024 · 33 citations
- Non-Coaxial Event-guided Motion Deblurring with Spatial AlignmentHoonhee Cho, Yuhwan Jeong, Taewoo Kim, Kuk-Jin YoonICCV 2023 · 30 citations
- Scene Adaptive Sparse Transformer for Event-based Object DetectionYansong Peng, Hebei Li, Yueyi Zhang, Xiaoyan Sun et al.CVPR 2024 · 25 citations
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
Related papers
- Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionLin Zhu, Tengyu Long, Xiao Wang, Lizhi Wang et al.NeurIPS 2025 · 4 citations
- EvRT-DETR: Latent Space Adaptation of Image Detectors for Event-Based VisionDmitrii Torbunov, Yihui Ren, Animesh Ghose, Odera Dim et al.ICCV 2025 · 7 citations
- Learning to Detect Objects with a 1 Megapixel Event CameraEtienne Perot, Pierre de Tournemire, Davide Nitti, Jonathan Masci et al.NeurIPS 2020 · 381 citations
- Event-based Video Reconstruction Using TransformerWenming Weng, Yueyi Zhang, Zhiwei XiongICCV 2021 · 139 citations
- Data-Driven Feature Tracking for Event CamerasNico Messikommer, Carter Fang, Mathias Gehrig, Davide ScaramuzzaCVPR 2023
