Distilling Monocular Foundation Model for Fine-grained Depth Completion
Yingping Liang, Yutao Hu, Wenqi Shao, Ying Fu
摘要
Depth completion involves predicting dense depth maps from sparse LiDAR inputs. However, sparse depth annotations from sensors limit the availability of dense supervision, which is necessary for learning detailed geometric features. In this paper, we propose a two-stage knowledge distillation framework that leverages powerful monocular foundation models to provide dense supervision for depth completion. In the first stage, we introduce a pre-training strategy that generates diverse training data from natural images, which distills geometric knowledge to depth completion. Specifically, we simulate LiDAR scans by utilizing monocular depth and mesh reconstruction, thereby creating training data without requiring ground-truth depth. Besides, monocular depth estimation suffers from inherent scale ambiguity in real-world settings. To address this, in the second stage, we employ a scale- and shift-invariant loss (SSI Loss) to learn real-world scales when fine-tuning on real-world datasets. Our two-stage distillation framework enables depth completion models to harness the strengths of monocular foundation models. Experimental results demonstrate that models trained with our two-stage distillation framework achieve state-of-the-art performance, ranking first place on the KITTI benchmark. Code is available at https://github.com/Sharpiless/DMD3C
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Event-Driven Dynamic Scene Depth CompletionZhiqiang Yan, Jianhao Jiao, Zhengxue Wang, Gim Hee LeeNeurIPS 2025 · 被引用 12 次
- DuCos: Duality Constrained Depth Super-Resolution via Foundation ModelZhiqiang Yan, Zhengxue Wang, Haoye Dong, Jun Li 等ICCV 2025 · 被引用 3 次
- The Midas Touch for Metric DepthYu Ma, Zizhan Guo, Zuyi Xiong, Haoran Zhang 等CVPR 2026 · 被引用 2 次
- PacGDC: Label-Efficient Generalizable Depth Completion with Projection Ambiguity and ConsistencyHaotian Wang, Aoran Xiao, Xiaoqin Zhang, Meng Yang 等ICCV 2025 · 被引用 1 次
- CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road TopographyGasser Elazab, Frank Neuhaus, Tilman Koß, Malte Splietker 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper29
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai 等ICCV 2023 · 被引用 388 次
- CSPN++: Learning Context and Resource Aware Convolutional Spatial Propagation Networks for Depth CompletionXinjing Cheng, Peng Wang, Chenye Guan, Ruigang YangAAAI 2020 · 被引用 270 次
相关 Paper
- Weakly Supervised Monocular 3D Detection with a Single-View ImageXueying Jiang, Sheng Jin, Lewei Lu, Xiaoqin Zhang 等CVPR 2024
- Attention-Based Depth Distillation with 3D-Aware Positional Encoding for Monocular 3D Object DetectionZizhang Wu, Yunzhe Wu, Jian Pu, Xianzhi Li 等AAAI 2023 · 被引用 29 次
- DesNet: Decomposed Scale-Consistent Network for Unsupervised Depth CompletionZhiqiang Yan, Kun Wang, Xiang Li, Zhenyu Zhang 等AAAI 2023 · 被引用 46 次
- Exploiting Pseudo Labels in a Self-Supervised Learning Framework for Improved Monocular Depth EstimationAndra Petrovai, Sergiu NedevschiCVPR 2022 · 被引用 56 次
- Distilling Diffusion Models to Efficient 3D LiDAR Scene CompletionShengyuan Zhang, An Zhao, Ling Yang, Zejian Li 等ICCV 2025 · 被引用 1 次
