From-Ground-To-Objects: Coarse-to-Fine Self-supervised Monocular Depth Estimation of Dynamic Objects with Ground Contact Prior
Jaeho Moon, Juan Luis Gonzalez Bello, Byeongjun Kwon, Munchurl Kim
Abstract
Self-supervised monocular depth estimation (DE) is an approach to learning depth without costly depth ground truths. However, it often struggles with moving objects that violate the static scene assumption during training. To address this issue, we introduce a coarse-to-fine training strategy leveraging the ground contacting prior based on the observation that most moving objects in outdoor scenes contact the ground. In the coarse training stage, we exclude the objects in dynamic classes from the reprojection loss calculation to avoid inaccurate depth learning. To provide precise supervision on the depth of the objects, we present a novel Ground-contacting-prior Disparity Smoothness Loss (GDS-Loss) that encourages a DE network to align the depth of the objects with their ground-contacting points. Subsequently, in the fine training stage, we refine the DE network to learn the detailed depth of the objects from the reprojection loss, while ensuring accurate DE on the moving object regions by employing our regularization loss with a cost-volume-based weighting factor. Our overall coarse-to-fine training strategy can easily be integrated with existing DE methods without any modifications, significantly enhancing DE performance on challenging Cityscapes and KITTI datasets, especially in the moving object regions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 026ab087-e6c4-4872-b15b-993c9cb07c81Cited by top-tier papers8
- Jasmine: Harnessing Diffusion Prior for Self-supervised Depth EstimationJiyuan Wang, Chunyu Lin, Cheng Guan, Lang Nie et al.NeurIPS 2025 · 26 citations
- SC-Lane: Slope-Aware and Consistent Road Height Estimation Framework for 3D Lane DetectionChaesong Park, Eunbin Seo, Jihyeon Hwang, Jongwoo LimICCV 2025 · 2 citations
- Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact EstimationDaniel Jung, Kyoung Mu LeeCVPR 2026 · 1 citation
- One Look is Enough: Seamless Patchwise Refinement for Zero-Shot Monocular Depth Estimation on High-Resolution ImagesByeongjun Kwon, Munchurl KimICCV 2025 · 1 citation
- Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono FailLuca Bartolomei, Fabio Tosi, Matteo Poggi, Stefano MattocciaCVPR 2025
Builds on27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 397 citations
- HR-Depth: High Resolution Self-Supervised Monocular Depth EstimationXiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong et al.AAAI 2021 · 341 citations
Related papers
- Mining Supervision for Dynamic Regions in Self-Supervised Monocular Depth EstimationHoang Chuong Nguyen, Tianyu Wang, José M. Álvarez, Miaomiao LiuCVPR 2024 · 5 citations
- Learning Occlusion-aware Coarse-to-Fine Depth Map for Self-supervised Monocular Depth EstimationZhengming Zhou, Qiulei DongACM MM 2022 · 21 citations
- AdaDepth: Exploiting Inherent Scene Information for Self-Supervised Depth Estimation in Dynamic ScenesXuanang Gao, Xiongbin Wu, Zhiwei Ning, Runze Yang et al.AAAI 2026
- Self-Supervised Monocular Depth HintsJamie Watson, Michael Firman, Gabriel J. Brostow, Daniyar TurmukhambetovICCV 2019 · 287 citations
- Exploiting Pseudo Labels in a Self-Supervised Learning Framework for Improved Monocular Depth EstimationAndra Petrovai, Sergiu NedevschiCVPR 2022 · 56 citations
