Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition from Egocentric RGB Videos
Yilin Wen, Hao Pan, Lei Yang, Jia Pan, Taku Komura, Wenping Wang
Abstract
Understanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit temporal information for robust estimation. Noticing the different temporal granularity of and the semantic correlation between hand pose estimation and action recognition, we build a network hierarchy with two cascaded transformer encoders, where the first one exploits the short-term temporal cue for hand pose estimation, and the latter aggregates per-frame pose and object information over a longer time span to recognize the action. Our approach achieves competitive results on two first-person hand action benchmarks, namely FPHA and H2O. Extensive ablation studies verify our design choices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object InteractionsHao Xu, Haipeng Li, Yinqiao Wang, Shuaicheng Liu et al.CVPR 2024 · 12 citations
- DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseHuangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu et al.AAAI 2025 · 8 citations
- Touchscreen-based Hand Tracking for Remote Whiteboard InteractionXinshuang Liu, Yizhong Zhang, Xin TongUIST 2024 · 8 citations
- CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionYuhang Wen, Mengyuan Liu, Songtao Wu, Beichen DingNeurIPS 2024 · 7 citations
- SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View AdaptationYinqiao Wang, Hao Xu, Pheng-Ann Heng, Chi-Wing FuAAAI 2024 · 5 citations
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai et al.ICCV 2019 · 504 citations
Related papers
- Transformer-based Unified Recognition of Two Hands Manipulating ObjectsHoseong Cho, Chanwoo Kim, Jihyeon Kim, Seongyeong Lee et al.CVPR 2023
- Deformer: Dynamic Fusion Transformer for Robust Hand Pose EstimationQichen Fu, Xingyu Liu, Ran Xu, Juan Carlos Niebles et al.ICCV 2023 · 27 citations
- Prior-Aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose EstimationPengfei Ren, Jingyu Wang, Haifeng Sun, Qi Qi et al.ICCV 2025 · 1 citation
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 69 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
