HARDVS: Revisiting Human Activity Recognition with Dynamic Vision Sensors
Xiao Wang, Zongzhen Wu, Bo Jiang, Zhimin Bao, Lin Zhu, Guoqi Li, Yaowei Wang, Yonghong Tian
Abstract
The main streams of human activity recognition (HAR) algorithms are developed based on RGB cameras which usually suffer from illumination, fast motion, privacy preservation, and large energy consumption. Meanwhile, the biologically inspired event cameras attracted great interest due to their unique features, such as high dynamic range, dense temporal but sparse spatial resolution, low latency, low power, etc. As it is a newly arising sensor, even there is no realistic large-scale dataset for HAR. Considering its great practical value, in this paper, we propose a large-scale benchmark dataset to bridge this gap, termed HARDVS, which contains 300 categories and more than 100K event sequences. We evaluate and report the performance of multiple popular HAR algorithms, which provide extensive baselines for future works to compare. More importantly, we propose a novel spatial-temporal feature learning and fusion framework, termed ESTF, for event stream based human activity recognition. It first projects the event streams into spatial and temporal embeddings using StemNet, then, encodes and fuses the dual-view representations using Transformer networks. Finally, the dual features are concatenated and fed into a classification head for activity prediction. Extensive experiments on multiple datasets fully validated the effectiveness of our model. Both the dataset and source code will be released at https://github.com/Event-AHU/HARDVS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e29836fe-2365-490f-b358-c3659b5b3b7cCited by top-tier papers14
- Inherent Redundancy in Spiking Neural NetworksMan Yao, Jiakui Hu, Guangshe Zhao, Yaoyuan Wang et al.ICCV 2023 · 30 citations
- UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural NetworksYuanbin Qian, Shuhan Ye, Chong Wang, Xiaojie Cai et al.AAAI 2025 · 18 citations
- ExACT: Language-Guided Conceptual Reasoning and Uncertainty Estimation for Event-Based Action Recognition and MoreJiazhou Zhou, Xu Zheng, Yuanhuiyi Lyu, Lin WangCVPR 2024 · 15 citations
- Finding Visual Saliency in Continuous Spike StreamLin Zhu, Xianzhang Chen, Xiao Wang, Hua HuangAAAI 2024 · 8 citations
- SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action RecognitionJing Wang, Rui Zhao, Ruiqin Xiong, Xingtao Wang et al.ICCV 2025 · 2 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
Related papers
- SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action RecognitionRui Fan, Weidong Hao, Juntao Guan, Lai Rui et al.CVPR 2026 · 1 citation
- DSF-Net: Dynamic Sparse Fusion of Event-RGB via Spike-Triggered Attention for High-Speed DetectionDongyang Ma, Zhengyu Ma, Wei Zhang, Yonghong TianACM MM 2025 · 1 citation
- Event Stream-Based Visual Object Tracking: A High-Resolution Benchmark Dataset and A Novel BaselineXiao Wang, Shiao Wang, Chuanming Tang, Lin Zhu et al.CVPR 2024 · 48 citations
- Hybrid Spiking Vision Transformer for Object Detection with Event CamerasQi Xu, Jie Deng, Jiangrong Shen, Biwu Chen et al.ICML 2025
- When Person Re-Identification Meets Event Camera: A Benchmark Dataset and an Attribute-Guided Re-Identification FrameworkXiao Wang, Qian Zhu, Shujuan Wu, Bo Jiang et al.AAAI 2026 · 2 citations
