DME: Unveiling the Bias for Better Generalized Monocular Depth Estimation
Songsong Yu, Yifan Wang, Yunzhi Zhuge, Lijun Wang, Huchuan Lu
Abstract
This paper aims to design monocular depth estimation models with better generalization abilities. To this end, we have conducted quantitative analysis and discovered two important insights. First, the Simulation Correlation phenomenon, commonly seen in long-tailed classification problems, also exists in monocular depth estimation, indicating that the imbalanced depth distribution in training data may be the cause of limited generalization ability. Second, the imbalanced and long-tail distribution of depth values extends beyond the dataset scale, and also manifests within each individual image, further exacerbating the challenge of monocular depth estimation. Motivated by the above findings, we propose the Distance-aware Multi-Expert (DME) depth estimation model. Unlike prior methods that handle different depth range indiscriminately, DME adopts a divide-and-conquer philosophy where each expert is responsible for depth estimation of regions within a specific depth range. As such, the depth distribution seen by each expert is more uniform and can be more easily predicted. A pixel-level routing module is further designed and learned to stitch the prediction of all experts into the final depth map. Experiments show that DME achieves state-of-the-art performance on both NYU-Depth v2 and KITTI, and also delivers favorable zero-shot generalization capability on unseen datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e03d2e1c-bde0-4b0f-bb82-c83a1ae83059Cited by top-tier papers4
- RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal FusionGeonho Bang, Minjae Seong, Jisong Kim, Geunju Baek et al.ICCV 2025 · 6 citations
- A Simple Yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision TasksQi Bi, Jingjun Yi, Huimin Huang, Hao Zheng et al.ICCV 2025 · 3 citations
- Mono2Stereo: A Benchmark and Empirical Study for Stereo ConversionSongsong Yu, Yuxin Chen, Zhongang Qi, Zeke Xie et al.CVPR 2025
- GeoDepth: From Point-to-Depth to Plane-to-Depth Modeling for Self-Supervised Monocular Depth EstimationHaifeng Wu, Shuhang Gu, Lixin Duan, Wen LiCVPR 2025
Builds on14
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Balanced Meta-Softmax for Long-Tailed Visual RecognitionJiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma et al.NeurIPS 2020 · 861 citations
- Long-tailed Recognition by Routing Diverse Distribution-Aware ExpertsXudong Wang, Long Lian, Zhongqi Miao, Ziwei Liu et al.ICLR 2021 · 481 citations
- Neural Window Fully-connected CRFs for Monocular Depth EstimationWeihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu et al.CVPR 2022 · 320 citations
Related papers
- Beyond the limitation of monocular 3D detector via knowledge distillationYiran Yang, Dongshuo Yin, Xuee Rong, Xian Sun et al.ICCV 2023 · 4 citations
- Attention-Based Depth Distillation with 3D-Aware Positional Encoding for Monocular 3D Object DetectionZizhang Wu, Yunzhe Wu, Jian Pu, Xianzhi Li et al.AAAI 2023 · 29 citations
- UniDAC: Universal Metric Depth Estimation for Any CameraGirish Chandar Ganesan, Yuliang Guo, Liu Ren, Xiaoming LiuCVPR 2026 · 8 citations
- MoGDE: Boosting Mobile Monocular 3D Object Detection with Ground Depth EstimationYunsong Zhou, Quan Liu, Hongzi Zhu, Yunzhe Li et al.NeurIPS 2022 · 23 citations
- GEDepth: Ground Embedding for Monocular Depth EstimationXiaodong Yang, Zhuang Ma, Zhiyu Ji, Zhe RenICCV 2023 · 40 citations
