S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing
Ping He, Rong Xiao, Wanying Xu, Chenwei Tang, Shudong Huang, Huajin Tang
Abstract
Dynamic Vision Sensor (DVS) asynchronously records sparse events triggered by changes in pixel intensity, offering high temporal resolution and low latency. Existing frame-based methods process event data densely, violating its inherent sparsity and introducing computational redundancy. While asynchronous models preserve the event stream's native format, they often neglect spatial information, compromising their adaptability and efficiency. To address these limitations, we propose a Spatiotemporally Separated Sparse Network (S3Net) for efficient event stream encoding and learning. Specifically, we employ a learnable sparse encoding scheme to construct a voxel-structured representation that effectively extracts spatiotemporal relationships among event data. After that, we propose a dual-branch architecture to capture localized spatial dependencies and dynamic temporal patterns of event data. By explicitly decoupling spatial and temporal modeling, S3Net enables end-to-end asynchronous processing of variable-length event sequences, achieving both strong representational capacity and high computational efficiency. Experimental results on six event-based datasets demonstrate that S3Net achieves state-of-the-art performance. Compared to frame-based methods, it significantly reduces computational overhead and model complexity, while also outperforming existing asynchronous approaches in inference speed without compromising accuracy. Extensive experiments across six event-based datasets show that S3Net establishes new state-of-the-art performance. Our method reduces computational costs by 35% and model parameters by 27% compared to frame-based approaches, while delivering 1.58× faster inference than existing point-based methods at comparable accuracy levels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffe31f66-022b-4230-89b2-4c9cc444bf79Builds on11
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 427 citations
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen et al.NeurIPS 2024 · 412 citations
- Graph-Based Object Classification for Neuromorphic Vision SensingYin Bi, Aaron Chadha, Alhabib Abbas, Eirina Bourtsoulatze et al.ICCV 2019 · 195 citations
Related papers
- Discrete time convolution for fast event-based stereoKaixuan Zhang, Kaiwei Che, Jianguo Zhang, Jie Cheng et al.CVPR 2022 · 34 citations
- TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event CamerasHongwei Ren, Yue Zhou, Haotian Fu, Yulong Huang et al.ACM MM 2023 · 14 citations
- AEDNet: Asynchronous Event Denoising with Spatial-Temporal Correlation among Irregular DataHuachen Fang, Jinjian Wu, Leida Li, Junhui Hou et al.ACM MM 2022 · 26 citations
- Leveraging Asynchronous Spiking Neural Networks for Ultra Efficient Event-Based Visual ProcessingDingyi Zeng, Yuchen Wang, Honglin Cao, Wanlong Liu et al.AAAI 2025 · 2 citations
- Dual Memory Aggregation Network for Event-Based Object Detection with Learnable RepresentationDongsheng Wang, Xu Jia, Yang Zhang, Xinyu Zhang et al.AAAI 2023 · 22 citations
