DeCoTR: Enhancing Depth Completion with 2D and 3D Attentions
Yunxiao Shi, Manish Kumar Singh, Hong Cai, Fatih Porikli
Abstract
In this paper, we introduce a novel approach that har-nesses both 2D and 3D attentions to enable highly accurate depth completion without requiring iterative spatial propa-gations. Specifically, we first enhance a baseline convolutional depth completion model by applying attention to 2D features in the bottleneck and skip connections. This effectively improves the performance of this simple network and sets it on par with the latest, complex transformer-based models. Leveraging the initial depths and features from this network, we uplift the 2D features to form a 3D point cloud and construct a 3D point transformer to process it, allowing the model to explicitly learn and exploit 3D geometric features. In addition, we propose normalization techniques to process the point cloud, which improves learning and leads to better accuracy than directly using point transformers off the shelf. Furthermore, we incorporate global attention on downsampled point cloud features, which enables long-range context while still being computationally feasible. We evaluate our method, DeCoTr, on established depth Completion benchmarks, including NYU Depth V2 and KITTI, showcasing that it sets new state-of-the-art performance. We further conduct zero-shot evaluations on ScanNet and DDAD benchmarks and demonstrate that DeCoTR has su-perior generalizability compared to existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf0d066a-60d3-4e8e-9ed8-c36a72273fd4Cited by top-tier papers1
Ask how each one uses itBuilds on16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu et al.NeurIPS 2022 · 924 citations
- Depth Completion From Sparse LiDAR Data With Depth-Normal ConstraintsYan Xu, Xinge Zhu, Jianping Shi, Guofeng Zhang et al.ICCV 2019 · 249 citations
- Learning Joint 2D-3D Representations for Depth CompletionYun Chen, Bin Yang, Ming Liang, Raquel UrtasunICCV 2019 · 190 citations
- Dynamic Spatial Propagation Network for Depth CompletionYuankai Lin, Tao Cheng, Qi Zhong, Wending Zhou et al.AAAI 2022 · 155 citations
Related papers
- CompletionFormer: Depth Completion with Convolutions and Vision TransformersYoumin Zhang, Xianda Guo, Matteo Poggi, Zheng Zhu et al.CVPR 2023
- Aggregating Feature Point Cloud for Depth CompletionZhu Yu, Zehua Sheng, Zili Zhou, Lun Luo et al.ICCV 2023 · 42 citations
- Context and Geometry Aware Voxel Transformer for Semantic Scene CompletionZhu Yu, Runmin Zhang, Jiacheng Ying, Junchen Yu et al.NeurIPS 2024 · 73 citations
- GeoFormer: Learning Point Cloud Completion with Tri-Plane Integrated TransformerJinpeng Yu, Binbin Huang, Yuxuan Zhang, Huaxia Li et al.ACM MM 2024 · 14 citations
- PointCFormer: A Relation-Based Progressive Feature Extraction Network for Point Cloud CompletionYi Zhong, Weize Quan, Dong-Ming Yan, Jie Jiang et al.AAAI 2025 · 3 citations
