Realistic Full-Body Tracking from Sparse Observations via Joint-Level Modeling
Xiaozheng Zheng, Zhuo Su, Chao Wen, Zhou Xue, Xiaojie Jin
Abstract
To bridge the physical and virtual worlds for rapidly developed VR/AR applications, the ability to realistically drive 3D full-body avatars is of great significance. Although real-time body tracking with only the head-mounted displays (HMDs) and hand controllers is heavily under-constrained, a carefully designed end-to-end neural network is of great potential to solve the problem by learning from large-scale motion data. To this end, we propose a two-stage framework that can obtain accurate and smooth full-body motions with the three tracking signals of head and hands only. Our framework explicitly models the joint-level features in the first stage and utilizes them as spatiotemporal tokens for alternating spatial and temporal transformer blocks to capture joint-level correlations in the second stage. Furthermore, we design a set of loss terms to constrain the task of a high degree of freedom, such that we can exploit the potential of our joint-level modeling. With extensive experiments on the AMASS motion dataset and real-captured data, we validate the effectiveness of our designs and show our proposed method can achieve more accurate and smooth motion compared to existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae3e6fed-cbbe-4b12-8d12-3935c48c5a1aCited by top-tier papers25
- Ultra Inertial Poser: Scalable Motion Capture and Tracking from Sparse Inertial Sensors and Ultra-Wideband RangingRayan Armani, Changlin Qian, Jiaxi Jiang, Christian HolzSIGGRAPH 2024 · 29 citations
- Physical Non-inertial Poser (PNP): Modeling Non-inertial Effects in Sparse-inertial Human Motion CaptureXinyu Yi, Yuxiao Zhou, Feng XuSIGGRAPH 2024 · 25 citations
- Transformer IMU Calibrator: Dynamic On-body IMU Calibration for Inertial Motion CaptureChengxu Zuo, Jiawei Huang, Xiao Jiang, Yuan Yao et al.SIGGRAPH 2025 · 13 citations
- EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUsZhen Fan, Peng Dai, Zhuo Su, Xu Gao et al.AAAI 2025 · 13 citations
- Accurate and Steady Inertial Pose Estimation through Sequence Structure Learning and ModulationYinghao Wu, Chaoran Wang, Lu Yin, Shihui Guo et al.NeurIPS 2024 · 11 citations
Builds on16
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- PhysDiff: Physics-Guided Human Motion Diffusion ModelYe Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat et al.ICCV 2023 · 414 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 399 citations
Related papers
- HMD-NeMo: Online 3D Avatar Motion Generation From Sparse ObservationsSadegh Aliakbarian, Fatemeh Sadat Saleh, David Collier, Pashmina Cameron et al.ICCV 2023 · 28 citations
- Avatars Grow Legs: Generating Smooth Human Motion from Sparse Tracking Inputs with Diffusion ModelYuming Du, Robin Kips, Albert Pumarola, Sebastian Starke et al.CVPR 2023
- HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse ObservationsPeng Dai, Yang Zhang, Tao Liu, Zhen Fan et al.CVPR 2024
- EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual RealityHaojie Cheng, Shaun Jing Heng Ong, Shaoyu Cai, Aiden Tat Yang Koh et al.IEEE VR 2026 · 1 citation
- A Unified Diffusion Framework for Scene-aware Human Motion Estimation from Sparse SignalsJiangnan Tang, Jingya Wang, Kaiyang Ji, Lan Xu et al.CVPR 2024 · 6 citations
