Hybrid Spiking Vision Transformer for Object Detection with Event Cameras
Qi Xu, Jie Deng, Jiangrong Shen, Biwu Chen, Huajin Tang, Gang Pan
摘要
Event-based object detection has attracted increasing attention for its high temporal resolution, wide dynamic range, and asynchronous address-event representation. Leveraging these advantages, spiking neural networks (SNNs) have emerged as a promising approach, offering low energy consumption and rich spatiotemporal dynamics. To further enhance the performance of event-based object detection, this study proposes a novel hybrid spike vision Transformer (HsVT) model. The HsVT model integrates a spatial feature extraction module to capture local and global features, and a temporal feature extraction module to model time dependencies and long-term patterns in event sequences. This combination enables HsVT to capture spatiotemporal features, improving its capability in handling complex event-based object detection tasks. To support research in this area, we developed the Fall Detection dataset as a benchmark for event-based object detection tasks. The Fall DVS detection dataset protects facial privacy and reduces memory usage thanks to its event-based representation. Experimental results demonstrate that HsVT outperforms existing SNN methods and achieves competitive performance compared to ANN-based models, with fewer parameters and lower energy consumption.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- FlashCap: Millisecond-Accurate Human Motion Capture via Flashing LEDs and Event-Based VisionZekai Wu, Shuqi Fan, Mengyin Liu, Yuhua Luo 等CVPR 2026 · 被引用 2 次
- Local-Global Coupling Spiking Graph Transformer for Brain Disorders Diagnosis from Two PerspectivesGeng Zhang, Jiangrong Shen, Kaizhong Zheng, Liangjun Chen 等NeurIPS 2025 · 被引用 1 次
- EvReflection: Event-Driven Micro-Dynamics for Reflection RemovalJiaxiao Wang, Dachun Kai, Huyue Zhu, Quanquan Hu 等ICML 2026
- Robust Selective Activation with Randomized Temporal K-Winner-Take-All in Spiking Neural Networks for Continual LearningJiangrong Shen, Liang Zhao, Qi Xu, Yuqi Yang 等ICLR 2026
- Bio-Vision-Inspired Spiking Neural Networks for Object Detection with Event CamerasDongyang Ma, Zhengyu Ma, Yifan Huang, Chenlin Zhou 等ICML 2026
它引用的顶会 Paper11
- Learning to Detect Objects with a 1 Megapixel Event CameraEtienne Perot, Pierre de Tournemire, Davide Nitti, Jonathan Masci 等NeurIPS 2020 · 被引用 381 次
- Deep Directly-Trained Spiking Neural Networks for Object DetectionQiaoyi Su, Yuhong Chou, Yifan Hu, Jianing Li 等ICCV 2023 · 被引用 143 次
- AEGNN: Asynchronous Event-based Graph Neural NetworksSimon Schaefer, Daniel Gehrig, Davide ScaramuzzaCVPR 2022 · 被引用 135 次
- GET: Group Event Transformer for Event-Based VisionYansong Peng, Yueyi Zhang, Zhiwei Xiong, Xiaoyan Sun 等ICCV 2023 · 被引用 86 次
- From Chaos Comes Order: Ordering Event Representations for Object Recognition and DetectionNikola Zubic, Daniel Gehrig, Mathias Gehrig, Davide ScaramuzzaICCV 2023 · 被引用 71 次
相关 Paper
- Spiking Transformers for Event-based Single Object TrackingJiqing Zhang, Bo Dong, Haiwei Zhang, Jianchuan Ding 等CVPR 2022 · 被引用 171 次
- Efficient Event-Based Object Detection: A Hybrid Neural Network with Spatial and Temporal AttentionSoikat Hasan Ahmed, Jan Finkbeiner, Emre NeftciCVPR 2025
- CREST: An Efficient Conjointly-trained Spike-driven Framework for Event-based Object Detection Exploiting Spatiotemporal DynamicsRuixin Mao, Aoyu Shen, Lin Tang, Jun ZhouAAAI 2025
- EGSST: Event-based Graph Spatiotemporal Sensitive Transformer for Object DetectionSheng Wu, Hang Sheng, Hui Feng, Bo HuNeurIPS 2024 · 被引用 9 次
- Event-based Video Reconstruction Using TransformerWenming Weng, Yueyi Zhang, Zhiwei XiongICCV 2021 · 被引用 139 次
