Scalable Autoregressive Monocular Depth Estimation
Jinhong Wang, Jian Liu, Dongqi Tang, Weiqiang Wang, Wentong Li, Danny Chen, Jintai Chen, Jian Wu
Abstract
This paper shows that the autoregressive model is an effective and scalable monocular depth estimator. Our idea is simple: We tackle the monocular depth estimation (MDE) task with an autoregressive prediction paradigm, based on two core designs. First, our depth autoregressive model (DAR) treats the depth map of different resolutions as a set of tokens, and conducts the low-to-high resolution autoregressive objective with a patch-wise causal mask. Second, our DAR recursively discretizes the entire depth range into more compact intervals, and attains the coarse-to-fine granularity autoregressive objective in an ordinal-regression manner. By coupling these two autoregressive objectives, our DAR establishes new state-ofthe-art (SOTA) on KITTI and NYU Depth v2 by clear margins. Further, our scalable approach allows us to scale the model up to 2.0B and achieve the best RMSE of 1.799 on the KITTI dataset (5% improvement) compared to 1.896 by the current SOTA (Depth Anything). DAR further showcases zero-shot generalization ability on unseen datasets. These results suggest that DAR yields superior performance with an autoregressive prediction paradigm, providing a promising approach to equip modern autoregressive large models (e.g., GPT-4o) with depth estimation capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32bf6bc1-1e98-4074-a85e-3e3e16bf3c52Cited by top-tier papers3
- DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language NavigationZihao Xin, Wentong Li, Yixuan Jiang, Bin Wang et al.CVPR 2026 · 6 citations
- GoR: A Unified and Extensible Generative Framework for Ordinal RegressionHongxu Ma, Han Zhou, Kai Tian, Xuefeng Zhang et al.ICLR 2026
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image ReconstructionXu Zhang, Ruijie Quan, Wenguan Wang, Yi YangICLR 2026
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
Related papers
- Towards Zero-Shot Scale-Aware Monocular Depth EstimationVitor Guizilini, Igor Vasiljevic, Dian Chen, Rares Ambrus et al.ICCV 2023 · 129 citations
- MAMo: Leveraging Memory and Attention for Monocular Video Depth EstimationRajeev Yasarla, Hong Cai, Jisoo Jeong, Yunxiao Shi et al.ICCV 2023 · 31 citations
- R-MSFM: Recurrent Multi-Scale Feature Modulation for Monocular Depth EstimatingZhongkai Zhou, Xinnan Fan, Pengfei Shi, Yuanxue XinICCV 2021 · 150 citations
- Depth Pro: Sharp Monocular Metric Depth in Less Than a SecondAlexey Bochkovskiy, Amaël Delaunoy, Hugo Germain, Marcel Santos et al.ICLR 2025 · 15 citations
- Self-Supervised Monocular Trained Depth Estimation Using Self-Attention and Discrete Disparity VolumeAdrian Johnston, Gustavo CarneiroCVPR 2020
