Single-Stage 3D Geometry-Preserving Depth Estimation Model Training on Dataset Mixtures with Uncalibrated Stereo Data
Nikolay Patakin, Anna Vorontsova, Mikhail Artemyev, Anton Konushin
摘要
Nowadays, robotics, AR, and 3D modeling applications attract considerable attention to single-view depth estimation (SVDE) as it allows estimating scene geometry from a single RGB image. Recent works have demonstrated that the accuracy of an SVDE method hugely depends on the diversity and volume of the training data. However, RGB-D datasets obtained via depth capturing or 3D re-construction are typically small, synthetic datasets are not photorealistic enough, and all these datasets lack diversity. The large-scale and diverse data can be sourced from stereo images or stereo videos from the web. Typically being uncalibrated, stereo data provides disparities up to unknown shift (geometrically incomplete data), so stereo-trained SVDE methods cannot recover 3D geometry. It was recently shown that the distorted point clouds obtained with a stereo-trained SVDE method can be corrected with additional point cloud modules (PCM) separately trained on the geometrically complete data. On the contrary, we propose <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> , General-Purpose and Geometry-Preserving training scheme, and show that conventional SVDE models can learn correct shifts themselves without any post-processing, benefiting from using stereo data even in the geometry-preserving setting. Through experiments on dif-ferent dataset mixtures, we prove that <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> -trained mod-els outperform methods relying on PCM in both accuracy and speed, and report the state-of-the-art results in the general-purpose geometry-preserving SVDE. Moreover, we show that SVDE models can learn to predict geometrically correct depth even when geometrically complete data com-prises the minor part of the training set.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ChildPlay: A New Benchmark for Understanding Children's Gaze BehaviourSamy Tafasca, Anshul Gupta, Jean-Marc OdobezICCV 2023 · 被引用 41 次
- Robust Geometry-Preserving Depth Estimation Using Differentiable RenderingChi Zhang, Wei Yin, Gang Yu, Zhibin Wang 等ICCV 2023 · 被引用 7 次
- Test-Time Prompt Tuning for Zero-Shot Depth CompletionChanhwi Jeong, Inhwan Bae, Jin-Hwi Park, Hae-Gon JeonICCV 2025 · 被引用 3 次
它引用的顶会 Paper5
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Monocular Depth Estimation via Listwise Ranking Using the Plackett-Luce ModelJulian Lienen, Eyke Hüllermeier, Ralph Ewerth, Nils NommensenCVPR 2021
- Structure-Guided Ranking Loss for Single Image Depth PredictionKe Xian, Jianming Zhang, Oliver Wang, Long Mai 等CVPR 2020
- Learning To Recover 3D Scene Shape From a Single ImageWei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus 等CVPR 2021
相关 Paper
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai 等ICCV 2023 · 被引用 388 次
- SynDeMo: Synergistic Deep Feature Alignment for Joint Learning of Depth and Ego-MotionBehzad Bozorgtabar, Mohammad Saeed Rad, Dwarikanath Mahapatra, Jean-Philippe ThiranICCV 2019 · 被引用 44 次
- SM4Depth: Seamless Monocular Metric Depth Estimation across Multiple Cameras and Scenes by One ModelYihao Liu, Feng Xue, Anlong Ming, Mingshuai Zhao 等ACM MM 2024 · 被引用 2 次
- MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training SupervisionRuicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang 等CVPR 2025
- 3D generation on ImageNetIvan Skorokhodov, Aliaksandr Siarohin, Yinghao Xu, Jian Ren 等ICLR 2023
