DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
Duolikun Danier, Mehmet Aygün, Changjian Li, Hakan Bilen, Oisin Mac Aodha
Abstract
Large-scale pre-trained vision models are becoming increasingly prevalent, offering expressive and generalizable visual representations that benefit various downstream tasks. Recent studies on the emergent properties of these models have revealed their high-level geometric understanding, in particular in the context of depth perception. However, it remains unclear how depth perception arises in these models without explicit depth supervision provided during pre-training. To investigate this, we examine whether the monocular depth cues, similar to those used by the human visual system, emerge in these models. We introduce a new benchmark, DepthCues, designed to evaluate depth cue understanding, and present findings across 20 diverse and representative pre-trained vision models. Our analysis shows that human-like depth cues emerge in more recent larger models. We also explore enhancing depth perception in large vision models by fine-tuning on DepthCues, and find that even without dense depth supervision, this improves depth estimation. To support further research, our benchmark and evaluation code will be made publicly available for studying depth perception in vision models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fa1a467-acbd-4b4c-b455-2bb256525613Cited by top-tier papers2
- Unique Lives, Shared World: Learning from Single-Life VideosTengda Han, Sayna Ebrahimi, Dilara Gokay, Li Yang Ku et al.CVPR 2026 · 2 citations
- MVSAnywhere: Zero-Shot Multi-View StereoSergio Izquierdo, Mohamed Sayed, Michael Firman, Guillermo Garcia-Hernando et al.CVPR 2025
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- GePBench: Evaluating Fundamental Geometric Perception for Multimodal Large Language ModelsShangyu Xing, Changhao Xiang, Xinyu Liu, Zhangtai Wu et al.ICML 2026
- MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface ReconstructionZehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler et al.NeurIPS 2022 · 670 citations
- Cue3D: Quantifying the Role of Image Cues in Single-Image 3D GenerationXiang Li, Zirui Wang, Zixuan Huang, James M. RehgNeurIPS 2025 · 2 citations
- V2Depth: Monocular Depth Estimation via Feature-Level Virtual-View Simulation and RefinementZizhang Wu, Zhuozheng Li, Zhi-Gang Fan, Yunzhe Wu et al.ACM MM 2023 · 4 citations
- 3D-IDE: 3D Implicit Depth EmergentChushan Zhang, Ruihan Lu, Jinguang Tong, Yikai Wang et al.CVPR 2026 · 1 citation
