SQLdepth: Generalizable Self-Supervised Fine-Structured Monocular Depth Estimation
Youhong Wang, Yunji Liang, Hao Xu, Shaohui Jiao, Hongkai Yu
Abstract
Recently, self-supervised monocular depth estimation has gained popularity with numerous applications in autonomous driving and robotics. However, existing solutions primarily seek to estimate depth from immediate visual features, and struggle to recover fine-grained scene details with limited generalization. In this paper, we introduce SQLdepth, a novel approach that can effectively learn fine-grained scene structures from motion. In SQLdepth, we propose a novel Self Query Layer (SQL) to build a selfcost volume and infer depth from it, rather than inferring depth from feature maps. The self-cost volume implicitly captures the intrinsic geometry of the scene within a single frame. Each individual slice of the volume signifies the relative distances between points and objects within a latent space. Ultimately, this volume is compressed to the depth map via a novel decoding approach. Experimental results on KITTI and Cityscapes show that our method attains remarkable state-of-the-art performance (AbsRel = 0.082 on KITTI, 0.052 on KITTI with improved ground-truth and 0.106 on Cityscapes), achieves 9.9%, 5.5% and 4.5% error reduction from the previous best. In addition, our approach showcases reduced training complexity, computational efficiency, improved generalization, and the ability to recover fine-grained scene details. Moreover, the selfsupervised pre-trained and metric fine-tuned SQLdepth can surpass existing supervised methods by significant margins (AbsRel = 0.043, 14% error reduction). Code is available at https://github.com/hisfog/SQLdepth-Impl .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be43d449-3ffa-4534-ab3a-7d7a471e7c33Cited by top-tier papers8
- RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language DescriptionsZiyao Zeng, Yangchao Wu, Hyoungseob Park, Daniel Wang et al.NeurIPS 2024 · 26 citations
- WorDepth: Variational Language Prior for Monocular Depth EstimationZiyao Zeng, Daniel Wang, Fengyu Yang, Hyoungseob Park et al.CVPR 2024 · 20 citations
- ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy PredictionYi Feng, Yu Han, Xijing Zhang, Tanghui Li et al.AAAI 2025 · 8 citations
- Intrinsic Image Decomposition for Robust Self-supervised Monocular Depth Estimation on Reflective SurfacesWonhyeok Choi, Kyumin Hwang, Minwoo Choi, Kiljoon Han et al.AAAI 2025 · 3 citations
- Hybrid-Grained Feature Aggregation with Coarse-to-Fine Language Guidance for Self-Supervised Monocular Depth EstimationWenyao Zhang, Hongsi Liu, Bohan Li, Jiawei He et al.ICCV 2025 · 2 citations
Builds on20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 397 citations
- HR-Depth: High Resolution Self-Supervised Monocular Depth EstimationXiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong et al.AAAI 2021 · 341 citations
Related papers
- Fine-grained Semantics-aware Representation Enhancement for Self-supervised Monocular Depth EstimationHyunyoung Jung, Eunhyeok Park, Sungjoo YooICCV 2021 · 133 citations
- Seeing Depth Through Frequency and Motion: A Progressive Training Paradigm for Monocular Depth EstimationKe Li, Bolin Song, Hongbo LiuCVPR 2026
- R-MSFM: Recurrent Multi-Scale Feature Modulation for Monocular Depth EstimatingZhongkai Zhou, Xinnan Fan, Pengfei Shi, Yuanxue XinICCV 2021 · 150 citations
- AdaDepth: Exploiting Inherent Scene Information for Self-Supervised Depth Estimation in Dynamic ScenesXuanang Gao, Xiongbin Wu, Zhiwei Ning, Runze Yang et al.AAAI 2026
- MonoIndoor: Towards Good Practice of Self-Supervised Monocular Depth Estimation for Indoor EnvironmentsPan Ji, Runze Li, Bir Bhanu, Yi XuICCV 2021 · 82 citations
