SAME: Spatial-Aware Multimodal Egocentric Human Pose Estimation
Yurong Fu, Peng Dai, Yu Zhang, Yiqiang Feng, Yang Zhang, Haoqian Wang
Abstract
Egocentric human pose estimation (HPE) plays a crucial role in immersive applications such as virtual and augmented reality. However, existing methods relying on either visual or sparse inertial data alone often suffer from occlusion or illposed problems. In this work, we propose SAME, a novel spatial-aware multi-modal fusion framework combining the complementary signals from the stereo images and sparse IMUs for accurate and robust egocentric HPE. It adopts a two-stage network based on a dual coordinate frame to mitigate the coordinate inconsistencies among the stereo cameras and the IMUs. In the first stage, the IMU signals are transformed into the local frame and iteratively fused with the stereo images for estimating 3D poses in the local frame. In the second stage, the local poses are transformed into the global frame with the 6DOF head poses provided by the headmounted display's (HMD) SLAM algorithm and then temporally aggregated via a temporal Transformer network. Meanwhile, to achieve geometric and semantic alignment among multi-modal features, we present a depth-guided spatialaware deformable stereo attention network and a modalityaware Transformer decoder for cross-view and cross-modal feature fusion. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on the public EMHI multi-modal egocentric pose estimation benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25d0f800-d5f1-4b1c-a4e5-e89b6996655eBuilds on16
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
- xR-EgoPose: Egocentric 3D Human Pose From an HMD CameraDenis Tomè, Patrick Peluse, Lourdes Agapito, Hernán BadinoICCV 2019 · 140 citations
- Realistic Full-Body Tracking from Sparse Observations via Joint-Level ModelingXiaozheng Zheng, Zhuo Su, Chao Wen, Zhou Xue et al.ICCV 2023 · 57 citations
- EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUsZhen Fan, Peng Dai, Zhuo Su, Xu Gao et al.AAAI 2025 · 13 citations
Related papers
- EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual RealityHaojie Cheng, Shaun Jing Heng Ong, Shaoyu Cai, Aiden Tat Yang Koh et al.IEEE VR 2026 · 1 citation
- Scene-Aware Egocentric 3D Human Pose EstimationJian Wang, Diogo C. Luvizon, Weipeng Xu, Lingjie Liu et al.CVPR 2023
- Domain-Guided Spatio-Temporal Self-Attention for Egocentric 3D Pose EstimationJinman Park, Kimathi Kaai, Saad Hossain, Norikatsu Sumi et al.KDD 2023 · 9 citations
- VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention NetworkZepeng Yang, Junxuan Bai, Hao Li, Ju Dai et al.CVPR 2026 · 1 citation
- Head2Body: Body Pose Generation from Multi-Sensory Head-Mounted InputsMinh Tran, Hongda Mao, Qingshuang Chen, Yelin KimICCV 2025 · 1 citation
