How Do Neural Networks See Depth in Single Images?
Tom van Dijk, Guido de Croon
Abstract
Deep neural networks have lead to a breakthrough in depth estimation from single images. Recent work often focuses on the accuracy of the depth map, where an evaluation on a publicly available test set such as the KITTI vision benchmark is often the main result of the article. While such an evaluation shows how well neural networks can estimate depth, it does not show how they do this. To the best of our knowledge, no work currently exists that analyzes what these networks have learned. In this work we take the MonoDepth network by Godard et al. and investigate what visual cues it exploits for depth estimation. We find that the network ignores the apparent size of known obstacles in favor of their vertical position in the image. Using the vertical position requires the camera pose to be known; however we find that MonoDepth only partially corrects for changes in camera pitch and roll and that these influence the estimated depth towards obstacles. We further show that MonoDepth's use of the vertical image position allows it to estimate the distance towards arbitrary obstacles, even those not appearing in the training set, but that it requires a strong edge at the ground contact point of the object to do so. In future work we will investigate whether these observations also apply to other neural networks for monocular depth estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d5ca15f-c437-412c-a8a4-e3d52dfd0eabCited by top-tier papers34
- Geometry Uncertainty Projection Network for Monocular 3D Object DetectionYan Lu, Xinzhu Ma, Lei Yang, Tianzhu Zhang et al.ICCV 2021 · 294 citations
- BEVStereo: Enhancing Depth Estimation in Multi-View 3D Object Detection with Temporal StereoYinhao Li, Han Bao, Zheng Ge, Jinrong Yang et al.AAAI 2023 · 226 citations
- Multi-Frame Self-Supervised Depth with TransformersVitor Guizilini, Rares Ambrus, Dian Chen, Sergey Zakharov et al.CVPR 2022 · 95 citations
- Self-supervised Monocular Depth Estimation for All Day Images using Domain SeparationLina Liu, Xibin Song, Mengmeng Wang, Yong Liu et al.ICCV 2021 · 95 citations
- Excavating the Potential Capacity of Self-Supervised Monocular Depth EstimationRui Peng, Ronggang Wang, Yawen Lai, Luyang Tang et al.ICCV 2021 · 93 citations
Related papers
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 199 citations
- MonoDGP: Monocular 3D Object Detection with Decoupled-Query and Geometry-Error PriorsFanqi Pu, Yifan Wang, Jiru Deng, Wenming YangCVPR 2025
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo et al.ICCV 2023 · 175 citations
- VA-DepthNet: A Variational Approach to Single Image Depth PredictionCe Liu, Suryansh Kumar, Shuhang Gu, Radu Timofte et al.ICLR 2023 · 17 citations
- Crafting Monocular Cues and Velocity Guidance for Self-Supervised Multi-Frame Depth LearningXiaofeng Wang, Zheng Zhu, Guan Huang, Xu Chi et al.AAAI 2023 · 31 citations
