Patch-Wise Attention Network for Monocular Depth Estimation
Sihaeng Lee, Janghyeon Lee, Byungju Kim, Eojindl Yi, Junmo Kim
Abstract
In computer vision, monocular depth estimation is the problem of obtaining a high-quality depth map from a two-dimensional image. This map provides information on three-dimensional scene geometry, which is necessary for various applications in academia and industry, such as robotics and autonomous driving. Recent studies based on convolutional neural networks achieved impressive results for this task. However, most previous studies did not consider the relationships between the neighboring pixels in a local area of the scene. To overcome the drawbacks of existing methods, we propose a patch-wise attention method for focusing on each local area. After extracting patches from an input feature map, our module generates attention maps for each local patch, using two attention modules for each patch along the channel and spatial dimensions. Subsequently, the attention maps return to their initial positions and merge into one attention feature. Our method is straightforward but effective. The experimental results on two challenging datasets, KITTI and NYU Depth V2, demonstrate that the proposed method achieves significant performance. Furthermore, our method outperforms other state-of-the-art methods on the KITTI depth estimation benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Neural Window Fully-connected CRFs for Monocular Depth EstimationWeihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu et al.CVPR 2022 · 320 citations
- Towards Zero-Shot Scale-Aware Monocular Depth EstimationVitor Guizilini, Igor Vasiljevic, Dian Chen, Rares Ambrus et al.ICCV 2023 · 129 citations
- IEBins: Iterative Elastic Bins for Monocular Depth EstimationShuwei Shao, Zhongcai Pei, Xingming Wu, Zhong Liu et al.NeurIPS 2023 · 114 citations
- Multi-Frame Self-Supervised Depth with TransformersVitor Guizilini, Rares Ambrus, Dian Chen, Sergey Zakharov et al.CVPR 2022 · 95 citations
- NDDepth: Normal-Distance Assisted Monocular Depth EstimationShuwei Shao, Zhongcai Pei, Weihai Chen, Xingming Wu et al.ICCV 2023 · 76 citations
Builds on1
Related papers
- Trap Attention: Monocular Depth Estimation with Manual TrapsChao Ning, Hongping GanCVPR 2023
- DCDepth: Progressive Monocular Depth Estimation in Discrete Cosine DomainKun Wang, Zhiqiang Yan, Junkai Fan, Wanlu Zhu et al.NeurIPS 2024 · 29 citations
- ROIFormer: Semantic-Aware Region of Interest Transformer for Efficient Self-Supervised Monocular Depth EstimationDaitao Xing, Jinglin Shen, Chiuman Ho, Anthony TzesAAAI 2023 · 17 citations
- Unsupervised High-Resolution Depth Learning From Videos With Dual NetworksJunsheng Zhou, Yuwang Wang, Kaihuai Qin, Wenjun ZengICCV 2019 · 77 citations
- Spatial Correspondence With Generative Adversarial Network: Learning Depth From Monocular VideosZhenyao Wu, Xinyi Wu, Xiaoping Zhang, Song Wang et al.ICCV 2019 · 29 citations
