Tame a Wild Camera: In-the-Wild Monocular Camera Calibration
Shengjie Zhu, Abhinav Kumar, Masa Hu, Xiaoming Liu
摘要
3D sensing for monocular in-the-wild images, e.g., depth estimation and 3D object detection, has become increasingly important. However, the unknown intrinsic parameter hinders their development and deployment. Previous methods for the monocular camera calibration rely on specific 3D objects or strong geometry prior, such as using a checkerboard or imposing a Manhattan World assumption. This work instead calibrates intrinsic via exploiting the monocular 3D prior. Given an undistorted image as input, our method calibrates the complete 4 Degree-of-Freedom (DoF) intrinsic parameters. First, we show intrinsic is determined by the two well-studied monocular priors: monocular depthmap and surface normal map. However, this solution necessitates a low-bias and low-variance depth estimation. Alternatively, we introduce the incidence field, defined as the incidence rays between points in 3D space and pixels in the 2D imaging plane. We show that: 1) The incidence field is a pixel-wise parametrization of the intrinsic invariant to image cropping and resizing. 2) The incidence field is a learnable monocular 3D prior, determined pixel-wisely by up-to-sacle monocular depthmap and surface normal. With the estimated incidence field, a robust RANSAC algorithm recovers intrinsic. We show the effectiveness of our method through superior performance on synthetic and zero-shot testing datasets. Beyond calibration, we demonstrate downstream applications in image manipulation detection & restoration, uncalibrated two-view pose estimation, and 3D sensing. Codes/models are held here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 等NeurIPS 2024 · 被引用 412 次
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han 等NeurIPS 2025 · 被引用 159 次
- From Seeing to Doing: Bridging Reasoning and Decision for Robotic ManipulationYifu Yuan, Haiqin Cui, Yibin Chen, Zibin Dong 等ICLR 2026 · 被引用 41 次
- InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language ModelsNianchen Deng, Lixin Gu, Shenglong Ye, Yinan He 等ICLR 2026 · 被引用 32 次
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language ModelsPingyi Chen, Yujing Lou, Shen Cao, Jinhui Guo 等NeurIPS 2025 · 被引用 24 次
它引用的顶会 Paper22
- DeepV2D: Video to Depth with Differentiable Structure from MotionZachary Teed, Jia DengICLR 2020 · 被引用 314 次
- Estimating and Exploiting the Aleatoric Uncertainty in Surface Normal EstimationGwangbin Bae, Ignas Budvytis, Roberto CipollaICCV 2021 · 被引用 154 次
- CTRL-C: Camera calibration TRansformer with Line-ClassificationJinwoo Lee, Hyunsung Go, Hyunjoon Lee, Sunghyun Cho 等ICCV 2021 · 被引用 52 次
- Voxel-based 3D Detection and Reconstruction of Multiple Objects from a Single ImageFeng Liu, Xiaoming LiuNeurIPS 2021 · 被引用 43 次
- Proactive Image Manipulation DetectionVishal Asnani, Xi Yin, Tal Hassner, Sijia Liu 等CVPR 2022 · 被引用 39 次
相关 Paper
- DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map GenerationXiankang He, Guangkai Xu, Bo Zhang, Hao Chen 等AAAI 2025 · 被引用 10 次
- AnyCalib: On-Manifold Learning for Model-Agnostic Single-View Camera CalibrationJavier Tirado-Garín, Javier CiveraICCV 2025 · 被引用 9 次
- Boost 3D Reconstruction Using Diffusion-Based Monocular Camera CalibrationJunyuan Deng, Wei Yin, Xiaoyang Guo, Qian Zhang 等ICCV 2025 · 被引用 2 次
- Ray3D: ray-based 3D human pose estimation for monocular absolute 3D localizationYu Zhan, Fenghai Li, Renliang Weng, Wongun ChoiCVPR 2022 · 被引用 62 次
- UniK3D: Universal Camera Monocular 3D EstimationLuigi Piccinelli, Christos Sakaridis, Mattia Segù, Yung-Hsu Yang 等CVPR 2025
