DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation
Xiankang He, Guangkai Xu, Bo Zhang, Hao Chen, Ying Cui, Dongyan Guo
Abstract
Monocular camera calibration is a key precondition for numerous 3D vision applications. Despite considerable advancements, existing methods often hinge on specific assumptions and struggle to generalize across varied real-world scenarios, and the performance is limited by insufficient training data. Recently, diffusion models trained on expansive datasets have been confirmed to maintain the capability to generate diverse, high-quality images. This success suggests a strong potential of the models to effectively understand varied visual information. In this work, we leverage the comprehensive visual knowledge embedded in pre-trained diffusion models to enable more robust and accurate monocular camera intrinsic estimation. Specifically, we reformulate the problem of estimating the four degrees of freedom (4-DoF) of camera intrinsic parameters as a dense incident map generation task. The map details the angle of incidence for each pixel in the RGB image, and its format aligns well with the paradigm of diffusion models. The camera intrinsic then can be derived from the incident map with a simple non-learning RANSAC algorithm during inference. Moreover, to further enhance the performance, we jointly estimate a depth map to provide extra geometric information for the incident map estimation. Extensive experiments on multiple testing datasets demonstrates that our model achieves state-of-the-art performance, gaining up to a 40% reduction in prediction errors. Besides, the experiments also show that the precise camera intrinsic and depth maps estimated by our pipeline can greatly benefit practical applications such as 3D reconstruction from a single in-the-wild image.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec8e1f3b-f5f3-4a2e-bedb-3de3cf29926bCited by top-tier papers7
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and GenerationKang Liao, Size Wu, Zhonghua Wu, Linyi Jin et al.ICLR 2026 · 19 citations
- AnyCalib: On-Manifold Learning for Model-Agnostic Single-View Camera CalibrationJavier Tirado-Garín, Javier CiveraICCV 2025 · 9 citations
- Detect Anything 3D in the WildHanxue Zhang, Haoran Jiang, Qingsong Yao, Yanan Sun et al.ICCV 2025 · 6 citations
- GeoMotion: Rethinking Motion Segmentation via Latent 4D GeometryXiankang He, Peile Lin, Ying Cui, Dongyan Guo et al.CVPR 2026 · 2 citations
- Boost 3D Reconstruction Using Diffusion-Based Monocular Camera CalibrationJunyuan Deng, Wei Yin, Xiaoyang Guo, Qian Zhang et al.ICCV 2025 · 2 citations
Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
Related papers
- Tame a Wild Camera: In-the-Wild Monocular Camera CalibrationShengjie Zhu, Abhinav Kumar, Masa Hu, Xiaoming LiuNeurIPS 2023 · 47 citations
- The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth EstimationSaurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar et al.NeurIPS 2023 · 160 citations
- Geo4D: Leveraging Video Generators for Geometric 4D Scene ReconstructionZeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus et al.ICCV 2025 · 9 citations
- PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle AdjustmentJianyuan Wang, Christian Rupprecht, David NovotnýICCV 2023 · 158 citations
- AlignDiff: Learning Physically-Grounded Camera Alignment via DiffusionLiuyue Xie, Jiancong Guo, Ozan Cakmakci, Andre Araujo et al.ICCV 2025
