GP2C: Geometric Projection Parameter Consensus for Joint 3D Pose and Focal Length Estimation in the Wild
Alexander Grabner, Peter M. Roth, Vincent Lepetit
Abstract
We present a joint 3D pose and focal length estimation approach for object categories in the wild. In contrast to previous methods that predict 3D poses independently of the focal length or assume a constant focal length, we explicitly estimate and integrate the focal length into the 3D pose estimation. For this purpose, we combine deep learning techniques and geometric algorithms in a two-stage approach: First, we estimate an initial focal length and establish 2D-3D correspondences from a single RGB image using a deep network. Second, we recover 3D poses and refine the focal length by minimizing the reprojection error of the predicted correspondences. In this way, we exploit the geometric prior given by the focal length for 3D pose estimation. This results in two advantages: First, we achieve significantly improved 3D translation and 3D pose accuracy compared to existing methods. Second, our approach finds a geometric consensus between the individual projection parameters, which is required for precise 2D-3D alignment. We evaluate our proposed approach on three challenging real-world datasets (Pix3D, Comp, and Stanford) with different object categories and significantly outperform the state-of-the-art by up to 20% absolute in multiple different metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5698dae2-e854-44ce-9623-1f325cb2cf86Cited by top-tier papers4
- Tame a Wild Camera: In-the-Wild Monocular Camera CalibrationShengjie Zhu, Abhinav Kumar, Masa Hu, Xiaoming LiuNeurIPS 2023 · 47 citations
- Focal Length and Object Pose Estimation via Render and CompareGeorgy Ponimatkin, Yann Labbé, Bryan C. Russell, Mathieu Aubry et al.CVPR 2022 · 18 citations
- DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map GenerationXiankang He, Guangkai Xu, Bo Zhang, Hao Chen et al.AAAI 2025 · 10 citations
- ClusterVO: Clustering Moving Instances and Estimating Visual Odometry for Self and SurroundingsJiahui Huang, Sheng Yang, Tai-Jiang Mu, Shi-Min HuCVPR 2020
Related papers
- Back to the Feature: Learning Robust Camera Localization From Pixels To PosePaul-Edouard Sarlin, Ajaykumar Unagar, Måns Larsson, Hugo Germain et al.CVPR 2021
- Universal Features Guided Zero-Shot Category-Level Object Pose EstimationWentian Qu, Chenyu Meng, Heng Li, Jian Cheng et al.AAAI 2025
- Ray3D: ray-based 3D human pose estimation for monocular absolute 3D localizationYu Zhan, Fenghai Li, Renliang Weng, Wongun ChoiCVPR 2022 · 62 citations
- CPPF: Towards Robust Category-Level 9D Pose Estimation in the WildYang You, Ruoxi Shi, Weiming Wang, Cewu LuCVPR 2022 · 37 citations
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
