AlignDiff: Learning Physically-Grounded Camera Alignment via Diffusion
Liuyue Xie, Jiancong Guo, Ozan Cakmakci, Andre Araujo, László A. Jeni, Zhiheng Jia
Abstract
Accurate camera calibration is a fundamental task for 3D perception, especially when dealing with real-world, in-the-wild environments where complex optical distortions are common. Existing methods often rely on pre-rectified images or calibration patterns, which limits their applicability and flexibility. In this work, we introduce a novel framework that addresses these challenges by jointly modeling camera intrinsic and extrinsic parameters using a generic ray camera model. Unlike previous approaches, AlignDiff shifts focus from semantic to geometric features, enabling more accurate modeling of local distortions. We propose AlignDiff, a diffusion model conditioned on geometric priors, enabling the simultaneous estimation of camera distortions and scene geometry. To enhance distortion prediction, we incorporate edge-aware attention, focusing the model on geometric features around image edges, rather than semantic content. Furthermore, to enhance generalizability to real-world captures, we incorporate a large database of ray-traced lenses containing over three thousand samples. This database characterizes the distortion inherent in a diverse variety of lens forms. Our experiments demonstrate that the proposed method significantly reduces the angular error of estimated ray bundles by 8.2 degrees and overall calibration accuracy, outperforming existing approaches on challenging, real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98ae4293-1219-4e72-9b24-7e9d07e24a3cCited by top-tier papers1
Ask how each one uses itBuilds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone et al.ICCV 2021 · 686 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle AdjustmentJianyuan Wang, Christian Rupprecht, David NovotnýICCV 2023 · 158 citations
- DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape GenerationShentong Mo, Enze Xie, Ruihang Chu, Lanqing Hong et al.NeurIPS 2023 · 157 citations
Related papers
- Why Having 10, 000 Parameters in Your Camera Model Is Better Than TwelveThomas Schöps, Viktor Larsson, Marc Pollefeys, Torsten SattlerCVPR 2020
- Cameras as Rays: Pose Estimation via Ray DiffusionJason Y. Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang et al.ICLR 2024 · 126 citations
- Self-Calibrating Neural Radiance FieldsYoonwoo Jeong, Seokjun Ahn, Christopher B. Choy, Animashree Anandkumar et al.ICCV 2021 · 275 citations
- Self-Calibrating Gaussian Splatting for Large Field-of-View ReconstructionYouming Deng, Wenqi Xian, Guandao Yang, Leonidas J. Guibas et al.ICCV 2025 · 1 citation
- AnyCalib: On-Manifold Learning for Model-Agnostic Single-View Camera CalibrationJavier Tirado-Garín, Javier CiveraICCV 2025 · 9 citations
