Spatial-Related Sensors Matters: 3D Human Motion Reconstruction Assisted with Textual Semantics
Xueyuan Yang, Chao Yao, Xiaojuan Ban
Abstract
Leveraging wearable devices for motion reconstruction has emerged as an economical and viable technique. Certain methodologies employ sparse Inertial Measurement Units (IMUs) on the human body and harness data-driven strategies to model human poses. However, the reconstruction of motion based solely on sparse IMUs data is inherently fraught with ambiguity, a consequence of numerous identical IMU readings corresponding to different poses. In this paper, we explore the spatial importance of multiple sensors, supervised by text that describes specific actions. Specifically, uncertainty is introduced to derive weighted features for each IMU. We also design a Hierarchical Temporal Transformer (HTT) and apply contrastive learning to achieve precise temporal and feature alignment of sensor data with textual semantics. Experimental results demonstrate our proposed approach achieves significant improvements in multiple metrics compared to existing methods. Notably, with textual supervision, our method not only differentiates between ambiguous actions such as sitting and standing but also produces more precise and natural motion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94ac1c2d-fa2d-4b9b-be07-6af5a12bb422Cited by top-tier papers2
- FisherPoser: Human Motion Estimation from Sparse Observations with Hierarchical Region-Wise Fisher-Matrix Uncertainty ModelingSongpengcheng Xia, Qingyu Zhang, Zhuo Su, Jiarui Yang et al.CVPR 2026
- EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty ModelingSongpengcheng Xia, Yu Zhang, Zhuo Su, Xiaozheng Zheng et al.CVPR 2025
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- TransPose: real-time 3D human translation and pose estimation with six inertial sensorsXinyu Yi, Yuxiao Zhou, Feng XuSIGGRAPH 2021 · 200 citations
Related papers
- MoBind: Motion Binding for Fine-Grained IMU-Video Pose AlignmentDuc Duy Nguyen, Tat-Jun Chin, Minh HoaiCVPR 2026 · 1 citation
- HiPoser: 3D Human Pose Estimation with Hierarchical Shared Learning at Parts-Level Using Inertial Measurement UnitsGuorui Liao, Chunyuan Zheng, Li Cheng, Haoyu Xie et al.AAAI 2025
- ToF-IP: Time-of-Flight Enhanced Sparse Inertial Poser for Real-time Human Motion CaptureYuan Yao, Shifan Jiang, Yangqing Hou, Chengxu Zuo et al.NeurIPS 2025 · 2 citations
- CTIN: Robust Contextual Transformer Network for Inertial NavigationBingbing Rao, Ehsan Kazemi, Yifan Ding, Devu M. Shila et al.AAAI 2022 · 66 citations
- Ultra Diffusion Poser: Diffusion-Based Human Motion Tracking from Sparse Inertial Sensors and Ranging-based Between-sensor DistancesDominik Hollidt, Tommaso Bendinelli, Christian HolzCVPR 2026
