State Space Models for Event Cameras
Nikola Zubic, Mathias Gehrig, Davide Scaramuzza
Abstract
Today, state-of-the-art deep neural networks that process event-camera data first convert a temporal window of events into dense, grid-like input representations. As such, they exhibit poor generalizability when deployed at higher inference frequencies (i.e., smaller temporal windows) than the ones they were trained on. We address this challenge by introducing state-space models (SSMs) with learnable timescale parameters to event-based vision. This design adapts to varying frequencies without the need to retrain the network at different frequencies. Additionally, we in-vestigate two strategies to counteract aliasing effects when deploying the model at higher frequencies. We compre-hensively evaluate our approach against existing methods based on RNN and Transformer architectures across various benchmarks, including Gen1 and 1 Mpx event camera datasets. Our results demonstrate that SSM-based models train 33% faster and also exhibit minimal performance degradation when tested at higher frequencies than the training input. Traditional RNN and Transformer models exhibit performance drops of more than 20 mAp, with SSMs having a drop of 3.31 mAp, highlighting the effectiveness of SSMs in event-based vision tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58e22bed-ac66-43db-a605-ff6c16b06d55Cited by top-tier papers31
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- SMamba: Sparse Mamba for Event-based Object DetectionNan Yang, Yang Wang, Zhanwen Liu, Meng Li et al.AAAI 2025 · 17 citations
- FlexEvent: Towards Flexible Event-Frame Object Detection at Varying Operational FrequenciesDongyue Lu, Lingdong Kong, Gim Hee Lee, Camille Simon Chane et al.NeurIPS 2025 · 13 citations
- EGSST: Event-based Graph Spatiotemporal Sensitive Transformer for Object DetectionSheng Wu, Hang Sheng, Hui Feng, Bo HuNeurIPS 2024 · 9 citations
- Talk2Event: Grounded Understanding of Dynamic Scenes from Event CamerasLingdong Kong, Dongyue Lu, Alan Liang, Rong Li et al.NeurIPS 2025 · 7 citations
Builds on15
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra et al.NeurIPS 2020 · 1,100 citations
- On the Parameterization and Initialization of Diagonal State Space ModelsAlbert Gu, Karan Goel, Ankit Gupta, Christopher RéNeurIPS 2022 · 690 citations
Related papers
- PASS: Path-selective State Space Model for Event-based RecognitionJiazhou Zhou, Kanghao Chen, Lei Zhang, Lin WangNeurIPS 2025 · 1 citation
- Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionLin Zhu, Tengyu Long, Xiao Wang, Lizhi Wang et al.NeurIPS 2025 · 4 citations
- Recurrent Vision Transformers for Object Detection with Event CamerasMathias Gehrig, Davide ScaramuzzaCVPR 2023
- DERD-Net: Learning Depth from Event-based Ray DensitiesDiego de Oliveira Hitzges, Suman Ghosh, Guillermo GallegoNeurIPS 2025 · 6 citations
- OmniEvent: Unified Event Representation LearningWeiqi Yan, Chenlu Lin, Youbiao Wang, Zhipeng Cai et al.AAAI 2026
