Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh Reconstruction
Wenjia Wang, Yongtao Ge, Haiyi Mei, Zhongang Cai, Qingping Sun, Yanjun Wang, Chunhua Shen, Lei Yang, Taku Komura
Abstract
As it is hard to calibrate single-view RGB images in the wild, existing 3D human mesh reconstruction (3DHMR) methods either use a constant large focal length or estimate one based on the background environment context, which can not tackle the problem of the torso, limb, hand or face distortion caused by perspective camera projection when the camera is close to the human body. The naive focal length assumptions can harm this task with the incorrectly formulated projection matrices. To solve this, we propose Zolly, the first 3DHMR method focusing on perspective-distorted images. Our approach begins with analysing the reason for perspective distortion, which we find is mainly caused by the relative location of the human body to the camera center. We propose a new camera model and a novel 2D representation, termed distortion image, which describes the 2D dense distortion scale of the human body. We then estimate the distance from distortion scale features rather than environment context features. Afterwards, We integrate the distortion feature with image features to reconstruct the body mesh. To formulate the correct projection matrix and locate the human body position, we simultaneously use perspective and weak-perspective projection loss. Since existing datasets could not handle this task, we propose the first synthetic dataset PDHuman and extend two real-world datasets tailored for this task, all containing perspective-distorted human images. Extensive experiments show that Zolly outperforms existing state-of-the-art methods on both perspective-distorted datasets and the standard benchmark (3DPW). Code and dataset will be released at https://wenjiawang0312.github.io/projects/zolly/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4e462ce-aab1-4d72-96c6-12a3a6fc11cfCited by top-tier papers4
- PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular VideosYufei Zhang, Jeffrey O. Kephart, Zijun Cui, Qiang JiCVPR 2024 · 14 citations
- Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh RecoveryYongwei Nie, Mingxian Fan, Chengjiang Long, Qing Zhang et al.NeurIPS 2024 · 1 citation
- CoEvoer: Collaborative Evolution Transformer for Upper-Body Expressive Human Pose and Shape EstimationYuxiang Zhao, Wei Huang, Yujie Song, Liu Wang et al.AAAI 2026
- DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single ImageQingxuan Wu, Zhiyang Dou, Sirui Xu, Soshi Shimada et al.ICLR 2025
Builds on17
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 399 citations
- PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback LoopHongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang et al.ICCV 2021 · 376 citations
Related papers
- SPEC: Seeing People in the Wild with an Estimated CameraMuhammed Kocabas, Chun-Hao P. Huang, Joachim Tesch, Lea Müller et al.ICCV 2021 · 181 citations
- BLADE: Single-view Body Mesh Estimation through Accurate Depth EstimationShengze Wang, Jiefeng Li, Tianye Li, Ye Yuan et al.CVPR 2025
- Perspose: 3D Human Pose Estimation with Perspective Encoding and Perspective RotationXiaoyang Hao, Han LiICCV 2025 · 3 citations
- Reconstructing 3D Human Pose by Watching Humans in the MirrorQi Fang, Qing Shuai, Junting Dong, Hujun Bao et al.CVPR 2021
- Learning Perspective Undistortion of PortraitsYajie Zhao, Zeng Huang, Tianye Li, Weikai Chen et al.ICCV 2019 · 29 citations
