SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognition
Rui Fan, Weidong Hao, Juntao Guan, Lai Rui, Tong Wu, Fanhong Zeng, Lin Gu
摘要
Event cameras action recognition (EAR) offers compelling privacy-protecting and efficiency advantages, where temporal motion dynamics is of great importance. Existing spatiotemporal multi-view representation learning (SMVRL) methods for event-based object recognition (EOR) offer promising solutions by projecting -- events alone spatial axis and , yet are limited by its translation-variant spatial binning representation and naive early concatenation fusion architecture. This paper reexamines the key SMVRL design stages for EAR and propose: (i) a principled spatiotemporal multi-view representation through translation-invariant dense conversion of sparse events, (ii) a dual-branch, dynamic fusion architecture that models sample-wise complementarity between motion features from different views, and (iii) a bio-inspired temporal warping augmentation that mimics speed variability of real-world human actions. On three challenging EAR datasets of HARDVS, DailyDVS-200 and THU-EACT-50-CHL, we show +7.0%, +10.7%, and +10.2% Top-1 accuracy gains over existing SMVRL EOR method with surprising 30.1% reduced parameters and 35.7% lower computations, establishing our framework as a novel and powerful EAR paradigm. Code will be released once accepted.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 被引用 427 次
- Spike-driven TransformerMan Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan 等NeurIPS 2023 · 被引用 368 次
相关 Paper
- HARDVS: Revisiting Human Activity Recognition with Dynamic Vision SensorsXiao Wang, Zongzhen Wu, Bo Jiang, Zhimin Bao 等AAAI 2024 · 被引用 80 次
- Event-Guided Person Re-Identification via Sparse-Dense Complementary LearningChengzhi Cao, Xueyang Fu, Hongjian Liu, Yukun Huang 等CVPR 2023
- Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionLin Zhu, Tengyu Long, Xiao Wang, Lizhi Wang 等NeurIPS 2025 · 被引用 4 次
- TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event CamerasHongwei Ren, Yue Zhou, Haotian Fu, Yulong Huang 等ACM MM 2023 · 被引用 14 次
- Event-Based Motion Deblurring Using Task-Oriented 3D Gaussian Event RepresentationsShengdong Xue, Haoxiang Ma, Hao Chen, Zhen Yang 等CVPR 2026
