S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing
Ping He, Rong Xiao, Wanying Xu, Chenwei Tang, Shudong Huang, Huajin Tang
摘要
Dynamic Vision Sensor (DVS) asynchronously records sparse events triggered by changes in pixel intensity, offering high temporal resolution and low latency. Existing frame-based methods process event data densely, violating its inherent sparsity and introducing computational redundancy. While asynchronous models preserve the event stream's native format, they often neglect spatial information, compromising their adaptability and efficiency. To address these limitations, we propose a Spatiotemporally Separated Sparse Network (S3Net) for efficient event stream encoding and learning. Specifically, we employ a learnable sparse encoding scheme to construct a voxel-structured representation that effectively extracts spatiotemporal relationships among event data. After that, we propose a dual-branch architecture to capture localized spatial dependencies and dynamic temporal patterns of event data. By explicitly decoupling spatial and temporal modeling, S3Net enables end-to-end asynchronous processing of variable-length event sequences, achieving both strong representational capacity and high computational efficiency. Experimental results on six event-based datasets demonstrate that S3Net achieves state-of-the-art performance. Compared to frame-based methods, it significantly reduces computational overhead and model complexity, while also outperforming existing asynchronous approaches in inference speed without compromising accuracy. Extensive experiments across six event-based datasets show that S3Net establishes new state-of-the-art performance. Our method reduces computational costs by 35% and model parameters by 27% compared to frame-based approaches, while delivering 1.58× faster inference than existing point-based methods at comparable accuracy levels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 被引用 427 次
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen 等NeurIPS 2024 · 被引用 412 次
- Graph-Based Object Classification for Neuromorphic Vision SensingYin Bi, Aaron Chadha, Alhabib Abbas, Eirina Bourtsoulatze 等ICCV 2019 · 被引用 195 次
相关 Paper
- Discrete time convolution for fast event-based stereoKaixuan Zhang, Kaiwei Che, Jianguo Zhang, Jie Cheng 等CVPR 2022 · 被引用 34 次
- TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event CamerasHongwei Ren, Yue Zhou, Haotian Fu, Yulong Huang 等ACM MM 2023 · 被引用 14 次
- AEDNet: Asynchronous Event Denoising with Spatial-Temporal Correlation among Irregular DataHuachen Fang, Jinjian Wu, Leida Li, Junhui Hou 等ACM MM 2022 · 被引用 26 次
- Leveraging Asynchronous Spiking Neural Networks for Ultra Efficient Event-Based Visual ProcessingDingyi Zeng, Yuchen Wang, Honglin Cao, Wanlong Liu 等AAAI 2025 · 被引用 2 次
- Dual Memory Aggregation Network for Event-Based Object Detection with Learnable RepresentationDongsheng Wang, Xu Jia, Yang Zhang, Xinyu Zhang 等AAAI 2023 · 被引用 22 次
