FisherPoser: Human Motion Estimation from Sparse Observations with Hierarchical Region-Wise Fisher-Matrix Uncertainty Modeling
Songpengcheng Xia, Qingyu Zhang, Zhuo Su, Jiarui Yang, Zengyuan Lai, Qi Wu, Ling Pei
Abstract
Full-body motion estimation from sparse VR observations is an inherently under-constrained problem, with only three 6-DoF trackers (HMD and controllers) available to infer a full skeletal pose. To address this ambiguity, we introduce a probabilistic framework that models joint orientations as distributions on SO(3) using the Matrix Fisher distribution. Instead of predicting a single deterministic pose, our network outputs a distribution for each joint, whose mode and concentration directly quantify prediction uncertainty on the rotation manifold. This enables likelihoodbased training and principled uncertainty propagation. At the core of our model is a causal Transformer encoder that fuses sparse observations with motion history. We further propose region-wise tokens for the torso, arms, and legs, obtained via attention pooling over local joint features and semantic VR anchors. These tokens guide compact, per-region Fisher regression. To ensure kinematic coherence efficiently, we employ a limb refinement module, where each child joint's Fisher parameters are conditioned on its parent's distribution and the regional context, propagating pose and uncertainty hierarchically. Extensive experiments on standard sparse-VR benchmarks show that our approach achieves state-of-the-art performance, while providing well-calibrated joint-wise uncertainty.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d8ba100-4d41-4872-91c6-57ff2b983790Builds on37
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
- TransPose: real-time 3D human translation and pose estimation with six inertial sensorsXinyu Yi, Yuxiao Zhou, Feng XuSIGGRAPH 2021 · 200 citations
- Physical Inertial Poser (PIP): Physics-aware Real-time Human Motion Tracking from Sparse Inertial SensorsXinyu Yi, Yuxiao Zhou, Marc Habermann, Soshi Shimada et al.CVPR 2022 · 198 citations
Related papers
- FLAG: Flow-based 3D Avatar Generation from Sparse ObservationsSadegh Aliakbarian, Pashmina Cameron, Federica Bogo, Andrew W. Fitzgibbon et al.CVPR 2022
- EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty ModelingSongpengcheng Xia, Yu Zhang, Zhuo Su, Xiaozheng Zheng et al.CVPR 2025
- Hierarchical Kinematic Probability Distributions for 3D Human Shape and Pose Estimation from Images in the WildAkash Sengupta, Ignas Budvytis, Roberto CipollaICCV 2021 · 68 citations
- Realistic Full-Body Tracking from Sparse Observations via Joint-Level ModelingXiaozheng Zheng, Zhuo Su, Chao Wen, Zhou Xue et al.ICCV 2023 · 57 citations
- Probabilistic Inertial Poser (ProbIP): Uncertainty-Aware Human Motion Modeling from Sparse Inertial SensorsMin Kim, Younho Jeon, Sungho JoICCV 2025 · 3 citations
