Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition from Egocentric RGB Videos
Yilin Wen, Hao Pan, Lei Yang, Jia Pan, Taku Komura, Wenping Wang
摘要
Understanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit temporal information for robust estimation. Noticing the different temporal granularity of and the semantic correlation between hand pose estimation and action recognition, we build a network hierarchy with two cascaded transformer encoders, where the first one exploits the short-term temporal cue for hand pose estimation, and the latter aggregates per-frame pose and object information over a longer time span to recognize the action. Our approach achieves competitive results on two first-person hand action benchmarks, namely FPHA and H2O. Extensive ablation studies verify our design choices.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object InteractionsHao Xu, Haipeng Li, Yinqiao Wang, Shuaicheng Liu 等CVPR 2024 · 被引用 12 次
- DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseHuangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu 等AAAI 2025 · 被引用 8 次
- Touchscreen-based Hand Tracking for Remote Whiteboard InteractionXinshuang Liu, Yizhong Zhang, Xin TongUIST 2024 · 被引用 8 次
- CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionYuhang Wen, Mengyuan Liu, Songtao Wu, Beichen DingNeurIPS 2024 · 被引用 7 次
- SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View AdaptationYinqiao Wang, Hao Xu, Pheng-Ann Heng, Chi-Wing FuAAAI 2024 · 被引用 5 次
它引用的顶会 Paper20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai 等ICCV 2019 · 被引用 504 次
相关 Paper
- Transformer-based Unified Recognition of Two Hands Manipulating ObjectsHoseong Cho, Chanwoo Kim, Jihyeon Kim, Seongyeong Lee 等CVPR 2023
- Deformer: Dynamic Fusion Transformer for Robust Hand Pose EstimationQichen Fu, Xingyu Liu, Ran Xu, Juan Carlos Niebles 等ICCV 2023 · 被引用 27 次
- Prior-Aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose EstimationPengfei Ren, Jingyu Wang, Haifeng Sun, Qi Qi 等ICCV 2025 · 被引用 1 次
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 被引用 69 次
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang 等CVPR 2022 · 被引用 403 次
