Head2Body: Body Pose Generation from Multi-Sensory Head-Mounted Inputs
Minh Tran, Hongda Mao, Qingshuang Chen, Yelin Kim
Abstract
Generating body pose from head-mounted, egocentric inputs is essential for immersive VR/AR and assistive technologies, as it supports more natural interactions. However, the task is challenging due to limited visibility of body parts in first-person views and the sparseness of sensory data, with only a single device placed on the head. To address these challenges, we introduce Head2Body, a novel framework for body pose estimation that effectively combines head-IMU and egocentric visual data. First, we introduce a pretrained IMU encoder, trained on over 1,700 hours of Ego4D IMU data from head-mounted devices, to better capture detailed temporal motion cues given limited labeled egocentric pose data. For visual processing, we leverage large vision-language models (LVLMs) to segment body parts that appear sporadically in video frames to improve visual feature extraction. To better guide pose generation from sparse head-mounted signals, we incorporate a residual Vector Quantized Variational Autoencoder (VQ-VAE) to represent poses with discrete tokens, capturing high-frequency motion patterns and improving over direct continuous regression, which often lacks structure and temporal consistency. Our experiments demonstrate the effectiveness of the proposed approach, yielding 6-13% gains over state-of-the-art baselines on three datasets: AMASS, KinPoly, and EgoExo4D. By capturing subtle temporal dynamics and leveraging complementary sensory data, our approach advances accurate egocentric body pose estimation and sets a new benchmark for multi-modal, first-person motion tracking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
Related papers
- EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual RealityHaojie Cheng, Shaun Jing Heng Ong, Shaoyu Cai, Aiden Tat Yang Koh et al.IEEE VR 2026 · 1 citation
- HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse ObservationsPeng Dai, Yang Zhang, Tao Liu, Zhen Fan et al.CVPR 2024
- Estimating Ego-Body Pose from Doubly Sparse Egocentric Video DataSeunggeun Chi, Pin-Hao Huang, Enna Sachdeva, Hengbo Ma et al.NeurIPS 2024 · 9 citations
- Ego-Body Pose Estimation via Ego-Head Pose EstimationJiaman Li, C. Karen Liu, Jiajun WuCVPR 2023
- HMD-NeMo: Online 3D Avatar Motion Generation From Sparse ObservationsSadegh Aliakbarian, Fatemeh Sadat Saleh, David Collier, Pashmina Cameron et al.ICCV 2023 · 28 citations
