Rethinking Scale-Aware Temporal Encoding for Event-based Object Detection
Lin Zhu, Tengyu Long, Xiao Wang, Lizhi Wang, Hua Huang
Abstract
Event cameras provide asynchronous, low-latency, and high-dynamic-range visual signals, making them ideal for real-time perception tasks such as object detection. However, effectively modeling the temporal dynamics of event streams remains a core challenge. Most existing methods follow frame-based detection paradigms, applying temporal modules only at high-level features, which limits early-stage temporal modeling. Transformer-based approaches introduce global attention to capture long-range dependencies, but often add unnecessary complexity and overlook fine-grained temporal cues. In this paper, we propose a CNN-RNN hybrid framework that rethinks temporal modeling for event-based object detection. Our approach is based on two key insights: (1) introducing recurrent modules at lower spatial scales to preserve detailed temporal information where events are most dense, and (2) utilizing Decoupled Deformable-enhanced Recurrent Layers specifically designed according to the inherent motion characteristics of event cameras to extract multiple spatiotemporal features, and performing independent downsampling at multiple spatiotemporal scales to enable flexible, scale-aware representation learning. These multi-scale features are then fused via a feature pyramid network to produce robust detection outputs. Experiments on Gen1, 1 Mpx and eTram dataset demonstrate that our approach achieves superior accuracy over recent transformer-based models, highlighting the importance of precise temporal feature extraction in early stages. This work offers a new perspective on designing architectures for event-driven vision beyond attention-centric paradigms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15d20aa5-9142-4e19-ad03-448eeb4d2d1fBuilds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- Learning to Detect Objects with a 1 Megapixel Event CameraEtienne Perot, Pierre de Tournemire, Davide Nitti, Jonathan Masci et al.NeurIPS 2020 · 381 citations
- Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual UnderstandingZizhao Zhang, Han Zhang, Long Zhao, Ting Chen et al.AAAI 2022 · 216 citations
Related papers
- Recurrent Vision Transformers for Object Detection with Event CamerasMathias Gehrig, Davide ScaramuzzaCVPR 2023
- Event-based Video Reconstruction Using TransformerWenming Weng, Yueyi Zhang, Zhiwei XiongICCV 2021 · 139 citations
- EvRT-DETR: Latent Space Adaptation of Image Detectors for Event-Based VisionDmitrii Torbunov, Yihui Ren, Animesh Ghose, Odera Dim et al.ICCV 2025 · 7 citations
- Efficient Event-Based Object Detection: A Hybrid Neural Network with Spatial and Temporal AttentionSoikat Hasan Ahmed, Jan Finkbeiner, Emre NeftciCVPR 2025
- EGSST: Event-based Graph Spatiotemporal Sensitive Transformer for Object DetectionSheng Wu, Hang Sheng, Hui Feng, Bo HuNeurIPS 2024 · 9 citations
