Rethinking Scale-Aware Temporal Encoding for Event-based Object Detection
Lin Zhu, Tengyu Long, Xiao Wang, Lizhi Wang, Hua Huang
摘要
Event cameras provide asynchronous, low-latency, and high-dynamic-range visual signals, making them ideal for real-time perception tasks such as object detection. However, effectively modeling the temporal dynamics of event streams remains a core challenge. Most existing methods follow frame-based detection paradigms, applying temporal modules only at high-level features, which limits early-stage temporal modeling. Transformer-based approaches introduce global attention to capture long-range dependencies, but often add unnecessary complexity and overlook fine-grained temporal cues. In this paper, we propose a CNN-RNN hybrid framework that rethinks temporal modeling for event-based object detection. Our approach is based on two key insights: (1) introducing recurrent modules at lower spatial scales to preserve detailed temporal information where events are most dense, and (2) utilizing Decoupled Deformable-enhanced Recurrent Layers specifically designed according to the inherent motion characteristics of event cameras to extract multiple spatiotemporal features, and performing independent downsampling at multiple spatiotemporal scales to enable flexible, scale-aware representation learning. These multi-scale features are then fused via a feature pyramid network to produce robust detection outputs. Experiments on Gen1, 1 Mpx and eTram dataset demonstrate that our approach achieves superior accuracy over recent transformer-based models, highlighting the importance of precise temporal feature extraction in early stages. This work offers a new perspective on designing architectures for event-driven vision beyond attention-centric paradigms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- Learning to Detect Objects with a 1 Megapixel Event CameraEtienne Perot, Pierre de Tournemire, Davide Nitti, Jonathan Masci 等NeurIPS 2020 · 被引用 381 次
- Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual UnderstandingZizhao Zhang, Han Zhang, Long Zhao, Ting Chen 等AAAI 2022 · 被引用 216 次
相关 Paper
- Recurrent Vision Transformers for Object Detection with Event CamerasMathias Gehrig, Davide ScaramuzzaCVPR 2023
- Event-based Video Reconstruction Using TransformerWenming Weng, Yueyi Zhang, Zhiwei XiongICCV 2021 · 被引用 139 次
- EvRT-DETR: Latent Space Adaptation of Image Detectors for Event-Based VisionDmitrii Torbunov, Yihui Ren, Animesh Ghose, Odera Dim 等ICCV 2025 · 被引用 7 次
- Efficient Event-Based Object Detection: A Hybrid Neural Network with Spatial and Temporal AttentionSoikat Hasan Ahmed, Jan Finkbeiner, Emre NeftciCVPR 2025
- EGSST: Event-based Graph Spatiotemporal Sensitive Transformer for Object DetectionSheng Wu, Hang Sheng, Hui Feng, Bo HuNeurIPS 2024 · 被引用 9 次
