AdaBins: Depth Estimation Using Adaptive Bins
Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka
Abstract
We address the problem of estimating a high quality dense depth map from a single RGB input image. We start out with a baseline encoder-decoder convolutional neural network architecture and pose the question of how the global processing of information can help improve overall depth estimation. To this end, we propose a transformerbased architecture block that divides the depth range into bins whose center value is estimated adaptively per image. The final depth values are estimated as linear combinations of the bin centers. We call our new building block AdaBins. Our results show a decisive improvement over the state-ofthe-art on several popular depth datasets across all metrics. We also validate the effectiveness of the proposed block with an ablation study and provide the code and corresponding pre-trained weights of the new state-of-the-art model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers209
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Learning to Upsample by Learning to SampleWenze Liu, Hao Lu, Hongtao Fu, Zhiguo CaoICCV 2023 · 518 citations
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai et al.ICCV 2023 · 388 citations
Builds on3
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 487 citations
- DepthLab: Real-time 3D Interaction with Depth Maps for Mobile Augmented RealityRuofei Du, Eric Turner, Maksym Dzitsiuk, Luca Prasso et al.UIST 2020 · 145 citations
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
Related papers
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- GuideFormer: Transformers for Image Guided Depth CompletionKyeongha Rho, Jinsung Ha, Youngjung KimCVPR 2022 · 57 citations
- CompletionFormer: Depth Completion with Convolutions and Vision TransformersYoumin Zhang, Xianda Guo, Matteo Poggi, Zheng Zhu et al.CVPR 2023
- Robust Multimodal Depth Estimation using Transformer based Generative Adversarial NetworksMd Fahim Faysal Khan, Anusha Devulapally, Siddharth Advani, Vijaykrishnan NarayananACM MM 2022 · 6 citations
- Transformer-Based Attention Networks for Continuous Pixel-Wise PredictionGuanglei Yang, Hao Tang, Mingli Ding, Nicu Sebe et al.ICCV 2021 · 246 citations
