HMD-NeMo: Online 3D Avatar Motion Generation From Sparse Observations
Sadegh Aliakbarian, Fatemeh Sadat Saleh, David Collier, Pashmina Cameron, Darren Cosker
Abstract
Generating both plausible and accurate full body avatar motion is the key to the quality of immersive experiences in mixed reality scenarios. Head-Mounted Devices (HMDs) typically only provide a few input signals, such as head and hands 6-DoF. Recently, different approaches achieved impressive performance in generating full body motion given only head and hands signal. However, to the best of our knowledge, all existing approaches rely on full hand visibility. While this is the case when, e.g., using motion controllers, a considerable proportion of mixed reality experiences do not involve motion controllers and instead rely on egocentric hand tracking. This introduces the challenge of partial hand visibility owing to the restricted field of view of the HMD. In this paper, we propose the first unified approach, HMD-NeMo, that addresses plausible and accurate full body motion generation even when the hands may be only partially visible. HMD-NeMo is a lightweight neural network that predicts the full body motion in an online and real-time fashion. At the heart of HMD-NeMo is the spatiotemporal encoder with novel temporally adaptable mask tokens that encourage plausible motion in the absence of hand observations. We perform extensive analysis of the impact of different components in HMD-NeMo and introduce a new state-of-the-art on AMASS dataset through our evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5268d9a0-9f74-464f-86ba-5c325545e6d3Cited by top-tier papers14
- Physical Non-inertial Poser (PNP): Modeling Non-inertial Effects in Sparse-inertial Human Motion CaptureXinyu Yi, Yuxiao Zhou, Feng XuSIGGRAPH 2024 · 25 citations
- Improving Global Motion Estimation in Sparse IMU-based Motion Capture with PhysicsXinyu Yi, Shaohua Pan, Feng XuSIGGRAPH 2025 · 7 citations
- A Unified Diffusion Framework for Scene-aware Human Motion Estimation from Sparse SignalsJiangnan Tang, Jingya Wang, Kaiyang Ji, Lan Xu et al.CVPR 2024 · 6 citations
- Zero-shot Human Pose Estimation using Diffusion-based Inverse solversSahil Bhandary Karnoor, Romit Roy ChoudhuryICLR 2026 · 2 citations
- Egocentric Visibility-Aware Human Pose EstimationPeng Dai, Yu Zhang, Feng Yiqiang, ZhenFan Fan et al.CVPR 2026 · 1 citation
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Probabilistic Modeling for Human Mesh RecoveryNikos Kolotouros, Georgios Pavlakos, Dinesh Jayaraman, Kostas DaniilidisICCV 2021 · 201 citations
- TransPose: real-time 3D human translation and pose estimation with six inertial sensorsXinyu Yi, Yuxiao Zhou, Feng XuSIGGRAPH 2021 · 200 citations
- Physical Inertial Poser (PIP): Physics-aware Real-time Human Motion Tracking from Sparse Inertial SensorsXinyu Yi, Yuxiao Zhou, Marc Habermann, Soshi Shimada et al.CVPR 2022 · 198 citations
Related papers
- Realistic Full-Body Tracking from Sparse Observations via Joint-Level ModelingXiaozheng Zheng, Zhuo Su, Chao Wen, Zhou Xue et al.ICCV 2023 · 57 citations
- HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse ObservationsPeng Dai, Yang Zhang, Tao Liu, Zhen Fan et al.CVPR 2024
- FLAG: Flow-based 3D Avatar Generation from Sparse ObservationsSadegh Aliakbarian, Pashmina Cameron, Federica Bogo, Andrew W. Fitzgibbon et al.CVPR 2022
- Avatars Grow Legs: Generating Smooth Human Motion from Sparse Tracking Inputs with Diffusion ModelYuming Du, Robin Kips, Albert Pumarola, Sebastian Starke et al.CVPR 2023
- EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual RealityHaojie Cheng, Shaun Jing Heng Ong, Shaoyu Cai, Aiden Tat Yang Koh et al.IEEE VR 2026 · 1 citation
