SATPose: Improving Monocular 3D Pose Estimation with Spatial-aware Ground Tactility
Lishuang Zhan, Enting Ying, Jiabao Gan, Shihui Guo, Boyu Gao, Yipeng Qin
Abstract
Estimating 3D human poses from monocular images is an important research area with many practical applications. However, the depth ambiguity of 2D solutions limits their accuracy in actions where occlusion exits or where slight centroid shifts can result in significant 3D pose variations. In this paper, we introduce a novel multimodal approach to mitigate the depth ambiguity inherent in monocular solutions by integrating spatial-aware pressure information. We first establish a data collection system with a pressure mat and a monocular camera, and construct a large-scale multimodal human activity dataset comprising over 600,000 frames of motion data. Utilizing this dataset, we propose a pressure image reconstruction network to extract pressure priors from monocular images. Subsequently, we introduce a Transformer-based multimodal pose estimation network to combine pressure priors with monocular images, achieving a world mean per joint position error of 51.6mm, outperforming state-of-the-art methods. Extensive experiments demonstrate the effectiveness of our multimodal 3D human pose estimation method across various actions and joints, highlighting the significance of spatial-aware pressure in improving the accuracy of monocular-vision-based methods. Our dataset is available at: https://github.com/LishuangZhan/SATPose.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 01d177b5-a2c4-4784-b9f7-1b27bbc482f7Cited by top-tier papers1
Ask how each one uses itRelated papers
- Intelligent Carpet: Inferring 3D Human Pose From Tactile SignalsYiyue Luo, Yunzhu Li, Michael Foshey, Wan Shou et al.CVPR 2021
- PressTrack-HMR: Pressure-Based Top-Down Multi-Person Global Human Mesh RecoveryJiayue Yuan, Fangting Xie, Guangwen Ouyang, Changhai Ma et al.AAAI 2026
- MMVP: A Multimodal MoCap Dataset with Vision and Pressure SensorsHe Zhang, Shenghao Ren, Haolei Yuan, Jianhui Zhao et al.CVPR 2024 · 10 citations
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt et al.ICCV 2019 · 108 citations
- MotionPRO: Exploring the Role of Pressure in Human MoCap and BeyondShenghao Ren, Yi Lu, Jiayi Huang, Jiayi Zhao et al.CVPR 2025
