Trap Attention: Monocular Depth Estimation with Manual Traps
Chao Ning, Hongping Gan
Abstract
Predicting a high quality depth map from a single image is a challenging task, because it exists infinite possibility to project a 2D scene to the corresponding 3D scene. Recently, some studies introduced multi-head attention (MHA) modules to perform long-range interaction, which have shown significant progress in regressing the depth maps. The main functions of MHA can be loosely summarized to capture long-distance information and report the attention map by the relationship between pixels. However, due to the quadratic complexity of MHA, these methods can not leverage MHA to compute depth features in high resolution with an appropriate computational complexity. In this paper, we exploit a depth-wise convolution to obtain long-range information, and propose a novel trap attention, which sets some traps on the extended space for each pixel, and forms the attention mechanism by the feature retention ratio of convolution window, resulting in that the quadratic computational complexity can be converted to linear form. Then we build an encoder-decoder trap depth estimation network, which introduces a vision transformer as the encoder, and uses the trap attention to estimate the depth from single image in the decoder. Extensive experimental results demonstrate that our proposed network can outperform the state-of-the-art methods in monocular depth estimation on datasets NYU Depth-v2 and KITTI, with significantly reduced number of parameters. Code is available at: https://github.com/ICSResearch/TrapAttention .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b5e1bc1-061d-40bd-bebc-5a3a32630fc1Cited by top-tier papers8
- Mind The Edge: Refining Depth Edges in Sparsely-Supervised Monocular Depth EstimationLior Talker, Aviad Cohen, Erez Yosef, Alexandra Dana et al.CVPR 2024 · 6 citations
- MD2E: Modeling Depth-to-Edge Cues for Monocular Metric Depth EstimationChao Ning, Minghe Shen, Naoto YokoyaCVPR 2026
- Shading Meets Motion: Self-supervised Indoor 3D Reconstruction Via Simultaneous Shape-from-Shading and Structure-from-MotionGuoyu LuCVPR 2025
- Hyper-Depth: Hypergraph-Based Multi-Scale Representation Fusion for Monocular Depth EstimationLin Bie, Siqi Li, Yifan Feng, Yue GaoICCV 2025
- SOAP: Vision-Centric 3D Semantic Scene Completion with Scene-Adaptive Decoder and Occluded Region-Aware View ProjectionHyo-Jun Lee, Yeong Jun Koh, Hanul Kim, Hyunseop Kim et al.CVPR 2025
Builds on11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
Related papers
- Patch-Wise Attention Network for Monocular Depth EstimationSihaeng Lee, Janghyeon Lee, Byungju Kim, Eojindl Yi et al.AAAI 2021 · 84 citations
- Transformer-Based Attention Networks for Continuous Pixel-Wise PredictionGuanglei Yang, Hao Tang, Mingli Ding, Nicu Sebe et al.ICCV 2021 · 246 citations
- Neural Window Fully-connected CRFs for Monocular Depth EstimationWeihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu et al.CVPR 2022 · 320 citations
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 199 citations
- CompletionFormer: Depth Completion with Convolutions and Vision TransformersYoumin Zhang, Xianda Guo, Matteo Poggi, Zheng Zhu et al.CVPR 2023
