Humans as Checkerboards: Calibrating Camera Motion Scale for World-Coordinate Human Mesh Recovery
Fengyuan Yang, Kerui Gu, Ha Linh Nguyen, Tze Ho Elden Tse, Angela Yao
Abstract
Accurate camera motion estimation is essential for recovering global human motion in world coordinates from RGB video inputs. While SLAM is widely used for estimating camera trajectory and point cloud, monocular SLAM does so only up to an unknown scale factor. Previous works estimate the scale factor through optimization, but this is unreliable and time-consuming. This paper presents an optimization-free scale calibration framework, Human as Checkerboard (HAC). HAC explicitly leverages the human body predicted by human mesh recovery model as a calibration reference. Specifically, it innovatively uses the absolute depth of human-scene contact joints as references to calibrate the corresponding relative scene depth from SLAM. HAC benefits from geometric priors encoded in human mesh recovery models to estimate the SLAM scale and achieves precise global human motion estimation. Simple yet powerful, our method sets a new state-of-the-art performance for global human mesh estimation tasks. It reduces motion errors by 50 % over prior local-to-global methods while using less post-SLAM inference time than optimization-based methods. Our code is available at https://martayang.github.io/HAC/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ab322e5-d712-4970-a90a-871a85e0e2b7Cited by top-tier papers1
Ask how each one uses itBuilds on21
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
Related papers
- WHAM: Reconstructing World-Grounded Humans with Accurate 3D MotionSoyong Shin, Juyong Kim, Eni Halilaj, Michael J. BlackCVPR 2024 · 66 citations
- Synergistic Global-Space Camera and Human Reconstruction from VideosYizhou Zhao, Tuanfeng Yang Wang, Bhiksha Raj, Min Xu et al.CVPR 2024 · 2 citations
- MetricHMSR: Metric Human Mesh and Scene Recovery from Monocular ImagesChentao Song, He Zhang, Haolei Yuan, Haozhe Lin et al.CVPR 2026 · 5 citations
- Reconstructing People, Places, and CamerasLea Müller, Hongsuk Choi, Anthony Zhang, Brent Yi et al.CVPR 2025
- PC-HMR: Pose Calibration for 3D Human Mesh Recovery from 2D Images/VideosTianyu Luan, Yali Wang, Junhao Zhang, Zhe Wang et al.AAAI 2021 · 45 citations
