Spatial-Related Sensors Matters: 3D Human Motion Reconstruction Assisted with Textual Semantics
Xueyuan Yang, Chao Yao, Xiaojuan Ban
摘要
Leveraging wearable devices for motion reconstruction has emerged as an economical and viable technique. Certain methodologies employ sparse Inertial Measurement Units (IMUs) on the human body and harness data-driven strategies to model human poses. However, the reconstruction of motion based solely on sparse IMUs data is inherently fraught with ambiguity, a consequence of numerous identical IMU readings corresponding to different poses. In this paper, we explore the spatial importance of multiple sensors, supervised by text that describes specific actions. Specifically, uncertainty is introduced to derive weighted features for each IMU. We also design a Hierarchical Temporal Transformer (HTT) and apply contrastive learning to achieve precise temporal and feature alignment of sensor data with textual semantics. Experimental results demonstrate our proposed approach achieves significant improvements in multiple metrics compared to existing methods. Notably, with textual supervision, our method not only differentiates between ambiguous actions such as sitting and standing but also produces more precise and natural motion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FisherPoser: Human Motion Estimation from Sparse Observations with Hierarchical Region-Wise Fisher-Matrix Uncertainty ModelingSongpengcheng Xia, Qingyu Zhang, Zhuo Su, Jiarui Yang 等CVPR 2026
- EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty ModelingSongpengcheng Xia, Yu Zhang, Zhuo Su, Xiaozheng Zheng 等CVPR 2025
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang 等CVPR 2022 · 被引用 462 次
- TransPose: real-time 3D human translation and pose estimation with six inertial sensorsXinyu Yi, Yuxiao Zhou, Feng XuSIGGRAPH 2021 · 被引用 200 次
相关 Paper
- MoBind: Motion Binding for Fine-Grained IMU-Video Pose AlignmentDuc Duy Nguyen, Tat-Jun Chin, Minh HoaiCVPR 2026 · 被引用 1 次
- HiPoser: 3D Human Pose Estimation with Hierarchical Shared Learning at Parts-Level Using Inertial Measurement UnitsGuorui Liao, Chunyuan Zheng, Li Cheng, Haoyu Xie 等AAAI 2025
- ToF-IP: Time-of-Flight Enhanced Sparse Inertial Poser for Real-time Human Motion CaptureYuan Yao, Shifan Jiang, Yangqing Hou, Chengxu Zuo 等NeurIPS 2025 · 被引用 2 次
- CTIN: Robust Contextual Transformer Network for Inertial NavigationBingbing Rao, Ehsan Kazemi, Yifan Ding, Devu M. Shila 等AAAI 2022 · 被引用 66 次
- Ultra Diffusion Poser: Diffusion-Based Human Motion Tracking from Sparse Inertial Sensors and Ranging-based Between-sensor DistancesDominik Hollidt, Tommaso Bendinelli, Christian HolzCVPR 2026
