Single-Stage 3D Geometry-Preserving Depth Estimation Model Training on Dataset Mixtures with Uncalibrated Stereo Data
Nikolay Patakin, Anna Vorontsova, Mikhail Artemyev, Anton Konushin
Abstract
Nowadays, robotics, AR, and 3D modeling applications attract considerable attention to single-view depth estimation (SVDE) as it allows estimating scene geometry from a single RGB image. Recent works have demonstrated that the accuracy of an SVDE method hugely depends on the diversity and volume of the training data. However, RGB-D datasets obtained via depth capturing or 3D re-construction are typically small, synthetic datasets are not photorealistic enough, and all these datasets lack diversity. The large-scale and diverse data can be sourced from stereo images or stereo videos from the web. Typically being uncalibrated, stereo data provides disparities up to unknown shift (geometrically incomplete data), so stereo-trained SVDE methods cannot recover 3D geometry. It was recently shown that the distorted point clouds obtained with a stereo-trained SVDE method can be corrected with additional point cloud modules (PCM) separately trained on the geometrically complete data. On the contrary, we propose <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> , General-Purpose and Geometry-Preserving training scheme, and show that conventional SVDE models can learn correct shifts themselves without any post-processing, benefiting from using stereo data even in the geometry-preserving setting. Through experiments on dif-ferent dataset mixtures, we prove that <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> -trained mod-els outperform methods relying on PCM in both accuracy and speed, and report the state-of-the-art results in the general-purpose geometry-preserving SVDE. Moreover, we show that SVDE models can learn to predict geometrically correct depth even when geometrically complete data com-prises the minor part of the training set.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0eb57307-e299-4f2b-a317-e9021c634e9eCited by top-tier papers3
- ChildPlay: A New Benchmark for Understanding Children's Gaze BehaviourSamy Tafasca, Anshul Gupta, Jean-Marc OdobezICCV 2023 · 41 citations
- Robust Geometry-Preserving Depth Estimation Using Differentiable RenderingChi Zhang, Wei Yin, Gang Yu, Zhibin Wang et al.ICCV 2023 · 7 citations
- Test-Time Prompt Tuning for Zero-Shot Depth CompletionChanhwi Jeong, Inhwan Bae, Jin-Hwi Park, Hae-Gon JeonICCV 2025 · 3 citations
Builds on5
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Monocular Depth Estimation via Listwise Ranking Using the Plackett-Luce ModelJulian Lienen, Eyke Hüllermeier, Ralph Ewerth, Nils NommensenCVPR 2021
- Structure-Guided Ranking Loss for Single Image Depth PredictionKe Xian, Jianming Zhang, Oliver Wang, Long Mai et al.CVPR 2020
- Learning To Recover 3D Scene Shape From a Single ImageWei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus et al.CVPR 2021
Related papers
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai et al.ICCV 2023 · 388 citations
- SynDeMo: Synergistic Deep Feature Alignment for Joint Learning of Depth and Ego-MotionBehzad Bozorgtabar, Mohammad Saeed Rad, Dwarikanath Mahapatra, Jean-Philippe ThiranICCV 2019 · 44 citations
- SM4Depth: Seamless Monocular Metric Depth Estimation across Multiple Cameras and Scenes by One ModelYihao Liu, Feng Xue, Anlong Ming, Mingshuai Zhao et al.ACM MM 2024 · 2 citations
- MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training SupervisionRuicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang et al.CVPR 2025
- 3D generation on ImageNetIvan Skorokhodov, Aliaksandr Siarohin, Yinghao Xu, Jian Ren et al.ICLR 2023
