Adaptive Vision Transformer for Event-Based Human Pose Estimation
Nannan Yu, Tao Ma, Jiqing Zhang, Yuji Zhang, Qirui Bao, Xiaopeng Wei, Xin Yang
Abstract
Event-based human pose estimation has gained popularity due to the benefits of high temporal resolution and high dynamic range offered by event cameras. The inherent spatial sparsity of event data makes discarding less significant regions a straightforward and effective way to decrease the computation. However, implementing this operation in CNNs poses a challenge, as it disrupts the regularity of dense convolutional workload. In this paper, we propose an adaptive vision transformer, a novel efficient backbone for human pose estimation with event cameras. Specifically, we present two adaptive patch and token sampling approaches based on the characteristics of events, thereby reducing the computational load while still achieving comparable performance. Firstly, we design an adaptive patch sampling scheme to eliminate inactivity patches by assessing the entropy of the events before they are inputted into the transformer. Secondly, we further propose an adaptive token reduction strategy to selectively remove less informative tokens in transformer layers through a dynamic token pruning algorithm. To exploit event-based visual cues in human pose estimation tasks, we construct a large-scale frame-event-based dataset, dubbed Event Multi Movement HPE (EventMM HPE). The dataset provides annotation frequencies up to 240 Hz. Extensive experiments demonstrate that our proposed approach outperforms existing state-of-the-art methods in estimation accuracy. The source code and dataset are available at https://github.com/doublemanyu/Adaptive-Vision-Transformer-for-Event-Based-HPE.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b0dbd8a2-48ad-492b-bba2-b3b01f2d5b91Related papers
- Eye-based Emotion Recognition via Event-Driven Sparse TransformersZixuan Wan, Jiqing Zhang, Yushan Wang, Hu Lin et al.ACM MM 2025 · 2 citations
- Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose EstimationWenhao Li, Mengyuan Liu, Hong Liu, Pichao Wang et al.CVPR 2024
- Event-based Video Reconstruction Using TransformerWenming Weng, Yueyi Zhang, Zhiwei XiongICCV 2021 · 139 citations
- EGSST: Event-based Graph Spatiotemporal Sensitive Transformer for Object DetectionSheng Wu, Hang Sheng, Hui Feng, Bo HuNeurIPS 2024 · 9 citations
- GET: Group Event Transformer for Event-Based VisionYansong Peng, Yueyi Zhang, Zhiwei Xiong, Xiaoyan Sun et al.ICCV 2023 · 86 citations
