Deep geometry-aware camera self-calibration from video
Annika Hagemann, Moritz Knorr, Christoph Stiller
Abstract
Accurate intrinsic calibration is essential for camera-based 3D perception, yet, it typically requires targets of well-known geometry. Here, we propose a camera self-calibration approach that infers camera intrinsics during application, from monocular videos in the wild. We propose to explicitly model projection functions and multi-view geometry, while leveraging the capabilities of deep neural networks for feature extraction and matching. To achieve this, we build upon recent research on integrating bundle adjustment into deep learning models, and introduce a self-calibrating bundle adjustment layer. The self-calibrating bundle adjustment layer optimizes camera intrinsics through classical Gauß-Newton steps and can be adapted to different camera models without re-training. As a specific realization, we implemented this layer within the deep visual SLAM system DROID-SLAM, and show that the resulting model, DroidCalib, yields state-of-the-art calibration accuracy across multiple public datasets. Our results suggest that the model generalizes to unseen environments and different camera models, including significant lens distortion. Thereby, the approach enables performing 3D perception tasks without prior knowledge about the camera. Code is available at https://github.com/boschresearch/droidcalib.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 304918ef-cff8-4c97-ac9d-c6c7ab72fd6bCited by top-tier papers9
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World ModelingYang Zhou, Yifan Wang, Jianjun Zhou, Wenzheng Chang et al.ICLR 2026 · 58 citations
- AnyCalib: On-Manifold Learning for Model-Agnostic Single-View Camera CalibrationJavier Tirado-Garín, Javier CiveraICCV 2025 · 9 citations
- AETHER: Geometric-Aware Unified World ModelingHaoyi Zhu, Yifan Wang, Jianjun Zhou, Wenzheng Chang et al.ICCV 2025 · 9 citations
- OpenVO: Open-World Visual Odometry with Temporal Dynamics AwarenessPhuc Nguyen, Anh N Nhu, Ming C. LinCVPR 2026 · 2 citations
- MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction PriorsRiku Murai, Eric Dexheimer, Andrew J. DavisonCVPR 2025
Builds on11
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 397 citations
- DeepV2D: Video to Depth with Differentiable Structure from MotionZachary Teed, Jia DengICLR 2020 · 314 citations
- Self-Calibrating Neural Radiance FieldsYoonwoo Jeong, Seokjun Ahn, Christopher B. Choy, Animashree Anandkumar et al.ICCV 2021 · 275 citations
- Pixel-Perfect Structure-from-Motion with Featuremetric RefinementPhilipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, Marc PollefeysICCV 2021 · 266 citations
Related papers
- MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic VideosZhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang et al.CVPR 2025
- Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and CameraYuhua Chen, Cordelia Schmid, Cristian SminchisescuICCV 2019 · 265 citations
- Deep Patch Visual OdometryZachary Teed, Lahav Lipson, Jia DengNeurIPS 2023 · 323 citations
- DROID-SLAM in the WildMoyang Li, Zihan Zhu, Marc Pollefeys, Daniel BarathCVPR 2026 · 10 citations
- AlignDiff: Learning Physically-Grounded Camera Alignment via DiffusionLiuyue Xie, Jiancong Guo, Ozan Cakmakci, Andre Araujo et al.ICCV 2025
