SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognition
Rui Fan, Weidong Hao, Juntao Guan, Lai Rui, Tong Wu, Fanhong Zeng, Lin Gu
Abstract
Event cameras action recognition (EAR) offers compelling privacy-protecting and efficiency advantages, where temporal motion dynamics is of great importance. Existing spatiotemporal multi-view representation learning (SMVRL) methods for event-based object recognition (EOR) offer promising solutions by projecting -- events alone spatial axis and , yet are limited by its translation-variant spatial binning representation and naive early concatenation fusion architecture. This paper reexamines the key SMVRL design stages for EAR and propose: (i) a principled spatiotemporal multi-view representation through translation-invariant dense conversion of sparse events, (ii) a dual-branch, dynamic fusion architecture that models sample-wise complementarity between motion features from different views, and (iii) a bio-inspired temporal warping augmentation that mimics speed variability of real-world human actions. On three challenging EAR datasets of HARDVS, DailyDVS-200 and THU-EACT-50-CHL, we show +7.0%, +10.7%, and +10.2% Top-1 accuracy gains over existing SMVRL EOR method with surprising 30.1% reduced parameters and 35.7% lower computations, establishing our framework as a novel and powerful EAR paradigm. Code will be released once accepted.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f839f584-af9e-4071-8b80-80a68bf4b53eBuilds on19
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 427 citations
- Spike-driven TransformerMan Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan et al.NeurIPS 2023 · 368 citations
Related papers
- HARDVS: Revisiting Human Activity Recognition with Dynamic Vision SensorsXiao Wang, Zongzhen Wu, Bo Jiang, Zhimin Bao et al.AAAI 2024 · 80 citations
- Event-Guided Person Re-Identification via Sparse-Dense Complementary LearningChengzhi Cao, Xueyang Fu, Hongjian Liu, Yukun Huang et al.CVPR 2023
- Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionLin Zhu, Tengyu Long, Xiao Wang, Lizhi Wang et al.NeurIPS 2025 · 4 citations
- TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event CamerasHongwei Ren, Yue Zhou, Haotian Fu, Yulong Huang et al.ACM MM 2023 · 14 citations
- Event-Based Motion Deblurring Using Task-Oriented 3D Gaussian Event RepresentationsShengdong Xue, Haoxiang Ma, Hao Chen, Zhen Yang et al.CVPR 2026
