HARDVS: Revisiting Human Activity Recognition with Dynamic Vision Sensors
Xiao Wang, Zongzhen Wu, Bo Jiang, Zhimin Bao, Lin Zhu, Guoqi Li, Yaowei Wang, Yonghong Tian
摘要
The main streams of human activity recognition (HAR) algorithms are developed based on RGB cameras which usually suffer from illumination, fast motion, privacy preservation, and large energy consumption. Meanwhile, the biologically inspired event cameras attracted great interest due to their unique features, such as high dynamic range, dense temporal but sparse spatial resolution, low latency, low power, etc. As it is a newly arising sensor, even there is no realistic large-scale dataset for HAR. Considering its great practical value, in this paper, we propose a large-scale benchmark dataset to bridge this gap, termed HARDVS, which contains 300 categories and more than 100K event sequences. We evaluate and report the performance of multiple popular HAR algorithms, which provide extensive baselines for future works to compare. More importantly, we propose a novel spatial-temporal feature learning and fusion framework, termed ESTF, for event stream based human activity recognition. It first projects the event streams into spatial and temporal embeddings using StemNet, then, encodes and fuses the dual-view representations using Transformer networks. Finally, the dual features are concatenated and fed into a classification head for activity prediction. Extensive experiments on multiple datasets fully validated the effectiveness of our model. Both the dataset and source code will be released at https://github.com/Event-AHU/HARDVS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Inherent Redundancy in Spiking Neural NetworksMan Yao, Jiakui Hu, Guangshe Zhao, Yaoyuan Wang 等ICCV 2023 · 被引用 30 次
- UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural NetworksYuanbin Qian, Shuhan Ye, Chong Wang, Xiaojie Cai 等AAAI 2025 · 被引用 18 次
- ExACT: Language-Guided Conceptual Reasoning and Uncertainty Estimation for Event-Based Action Recognition and MoreJiazhou Zhou, Xu Zheng, Yuanhuiyi Lyu, Lin WangCVPR 2024 · 被引用 15 次
- Finding Visual Saliency in Continuous Spike StreamLin Zhu, Xianzhang Chen, Xiao Wang, Hua HuangAAAI 2024 · 被引用 8 次
- SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action RecognitionJing Wang, Rui Zhao, Ruiqin Xiong, Xingtao Wang 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
相关 Paper
- SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action RecognitionRui Fan, Weidong Hao, Juntao Guan, Lai Rui 等CVPR 2026 · 被引用 1 次
- DSF-Net: Dynamic Sparse Fusion of Event-RGB via Spike-Triggered Attention for High-Speed DetectionDongyang Ma, Zhengyu Ma, Wei Zhang, Yonghong TianACM MM 2025 · 被引用 1 次
- Event Stream-Based Visual Object Tracking: A High-Resolution Benchmark Dataset and A Novel BaselineXiao Wang, Shiao Wang, Chuanming Tang, Lin Zhu 等CVPR 2024 · 被引用 48 次
- Hybrid Spiking Vision Transformer for Object Detection with Event CamerasQi Xu, Jie Deng, Jiangrong Shen, Biwu Chen 等ICML 2025
- When Person Re-Identification Meets Event Camera: A Benchmark Dataset and an Attribute-Guided Re-Identification FrameworkXiao Wang, Qian Zhu, Shujuan Wu, Bo Jiang 等AAAI 2026 · 被引用 2 次
