Domain-Guided Spatio-Temporal Self-Attention for Egocentric 3D Pose Estimation
Jinman Park, Kimathi Kaai, Saad Hossain, Norikatsu Sumi, Sirisha Rambhatla, Paul W. Fieguth
Abstract
Vision-based ego-centric 3D human pose estimation (ego-HPE) is essential to support critical applications of xR-technologies. However, severe self-occlusions and strong distortion introduced by the fish-eye view from the head mounted camera, make ego-HPE extremely challenging. To address these challenges, we propose a domain-guided spatio-temporal transformer model that leverages information specific to ego-views. Powered by this domain-guided transformer, we build Egocentric Spatio-Temporal Self-Attention Network (Ego-STAN), which uses 2D image representations and spatio-temporal attention to address both distortions and self-occlusions in ego-HPE. Additionally, we introduce a spatial concept called feature map tokens (FMT) which endows Ego-STAN with the ability to draw complex spatio-temporal information encoded in ego-centric videos. Our quantitative evaluation on the contemporary xR-EgoPose dataset, achieves a 38.2% improvement on the highest error joints against the SOTA ego-HPE model, while accomplishing a 22% decrease in the number of parameters. Finally, we also demonstrate the generalization capabilities of our model to real-world HPE tasks beyond ego-views achieving 7.7% improvement on 2D human pose estimation with the Human3.6M dataset. Our code is also made available at: https://github.com/jmpark0808/Ego-STAN
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e9f2a7fe-bd1a-438d-b3d0-c2ac2b3e78d4Cited by top-tier papers8
- Egocentric Whole-Body Motion Capture with FisheyeViT and Diffusion-Based Motion RefinementJian Wang, Zhe Cao, Diogo C. Luvizon, Lingjie Liu et al.CVPR 2024 · 18 citations
- EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUsZhen Fan, Peng Dai, Zhuo Su, Xu Gao et al.AAAI 2025 · 13 citations
- Bring Your Rear Cameras for Egocentric 3D Human Pose EstimationHiroyasu Akada, Jian Wang, Vladislav Golyanik, Christian TheobaltICCV 2025 · 10 citations
- Ego3DT: Tracking Every 3D Object in Ego-centric VideosShengyu Hao, Wenhao Chai, Zhonghan Zhao, Meiqi Sun et al.ACM MM 2024 · 6 citations
- UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and GenerationChaitanya Patel, Hiroki Nakamura, Yuta Kyuragi, Kazuki Kozuka et al.ICCV 2025 · 3 citations
Related papers
- 3D Human Pose Perception from Egocentric Stereo VideosHiroyasu Akada, Jian Wang, Vladislav Golyanik, Christian TheobaltCVPR 2024
- xR-EgoPose: Egocentric 3D Human Pose From an HMD CameraDenis Tomè, Patrick Peluse, Lourdes Agapito, Hernán BadinoICCV 2019 · 140 citations
- Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric VisionTianma Shen, Aditya Puranik, James Vong, Vrushabh Abhijit Deogirikar et al.ICCV 2025 · 1 citation
- Scene-Aware Egocentric 3D Human Pose EstimationJian Wang, Diogo C. Luvizon, Weipeng Xu, Lingjie Liu et al.CVPR 2023
- Estimating Egocentric 3D Human Pose in Global SpaceJian Wang, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar et al.ICCV 2021 · 78 citations
