Depth Anything V2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao
Abstract
This work presents Depth Anything V2. Without pursuing fancy techniques, we aim to reveal crucial findings to pave the way towards building a powerful monocular depth estimation model. Notably, compared with V1, this version produces much finer and more robust depth predictions through three key practices: 1) replacing all labeled real images with synthetic images, 2) scaling up the capacity of our teacher model, and 3) teaching student models via the bridge of large-scale pseudo-labeled real images. Compared with the latest models built on Stable Diffusion, our models are significantly more efficient (more than 10x faster) and more accurate. We offer models of different scales (ranging from 25M to 1.3B params) to support extensive scenarios. Benefiting from their strong generalization capability, we fine-tune them with metric depth labels to obtain our metric depth models. In addition to our models, considering the limited diversity and frequent noise in current test sets, we construct a versatile evaluation benchmark with precise annotations and diverse scenes to facilitate future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a214d4f2-a6d9-44a8-b959-aede37dfb265Cited by top-tier papers529
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen et al.ICLR 2026 · 720 citations
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang et al.ICLR 2026 · 318 citations
- MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsRuicheng Wang, Sicheng Xu, Yue Dong, Yu Deng et al.NeurIPS 2025 · 308 citations
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang et al.NeurIPS 2025 · 244 citations
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han et al.NeurIPS 2025 · 159 citations
Builds on49
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Depth Any Video with Scalable Synthetic DataHonghui Yang, Di Huang, Wei Yin, Chunhua Shen et al.ICLR 2025
- FiffDepth: Feed-Forward Transformation of Diffusion-Based Generators for Detailed Depth EstimationYunpeng Bai, Qixing HuangICCV 2025 · 5 citations
- Repurposing Diffusion-Based Image Generators for Monocular Depth EstimationBingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger et al.CVPR 2024
- Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono FailLuca Bartolomei, Fabio Tosi, Matteo Poggi, Stefano MattocciaCVPR 2025
