Sparkle: A Robust and Versatile Representation for Point Cloud-based Human Motion Capture
Yiming Ren, Yujing Sun, Aoru Xue, Kwok-Yan Lam, Yuexin Ma
Abstract
Point cloud-based motion capture leverages rich spatial geometry and privacy-preserving sensing, but learning robust representations from noisy, unstructured point clouds remains challenging. Existing approaches face a struggle trade-off between point-based methods (geometrically detailed but noisy) and skeleton-based ones (robust but oversimplified). We address the fundamental challenge: how to construct an effective representation for human motion capture that can balance expressiveness and robustness. In this paper, we propose Sparkle, a structured representation unifying skeletal joints and surface anchors with explicit kinematic-geometric factorization. Our framework, SparkleMotion, learns this representation through hierarchical modules embedding geometric continuity and kinematic constraints. By explicitly disentangling internal kinematic structure from external surface geometry, SparkleMotion achieves state-of-the-art performance not only in accuracy but crucially in robustness and generalization under severe domain shifts, noise, and occlusion. Extensive experiments demonstrate our superiority across diverse sensor types and challenging real-world scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- SoftGroup for 3D Instance Segmentation on Point CloudsThang Vu, Kookhoi Kim, Tung Minh Luu, Thanh Xuan Nguyen et al.CVPR 2022 · 251 citations
- TransPose: real-time 3D human translation and pose estimation with six inertial sensorsXinyu Yi, Yuxiao Zhou, Feng XuSIGGRAPH 2021 · 200 citations
- Physical Inertial Poser (PIP): Physics-aware Real-time Human Motion Tracking from Sparse Inertial SensorsXinyu Yi, Yuxiao Zhou, Marc Habermann, Soshi Shimada et al.CVPR 2022 · 198 citations
- WHAM: Reconstructing World-Grounded Humans with Accurate 3D MotionSoyong Shin, Juyong Kim, Eni Halilaj, Michael J. BlackCVPR 2024 · 66 citations
Related papers
- SPEAL: Skeletal Prior Embedded Attention Learning for Cross-Source Point Cloud RegistrationKezheng Xiong, Maoji Zheng, Qingshan Xu, Chenglu Wen et al.AAAI 2024 · 24 citations
- Glimpse: Geometry Learning of Multi-scale Structural Priors for 3D Pose EstimationZhenhua TANG, Jihua Peng, Yanbin Hao, Qiguang Miao et al.ICML 2026
- Geo-CF2Net: Geometry-Prior Cross-Frequency Interactive Fusion Network for 3D Human Action RecognitionZhaoyu Chen, Qian Huang, Xing Li, Yunfei Zhang et al.ACM MM 2025
- SOMA: Solving Optical Marker-Based MoCap AutomaticallyNima Ghorbani, Michael J. BlackICCV 2021 · 48 citations
- VoteHMR: Occlusion-Aware Voting Network for Robust 3D Human Mesh Recovery from Partial Point CloudsGuanze Liu, Yu Rong, Lu ShengACM MM 2021 · 27 citations
