CompletionFormer: Depth Completion with Convolutions and Vision Transformers
Youmin Zhang, Xianda Guo, Matteo Poggi, Zheng Zhu, Guan Huang, Stefano Mattoccia
Abstract
Given sparse depths and the corresponding RGB images, depth completion aims at spatially propagating the sparse measurements throughout the whole image to get a dense depth prediction. Despite the tremendous progress of deeplearning-based depth completion methods, the locality of the convolutional layer or graph model makes it hard for the network to model the long-range relationship between pixels. While recent fully Transformer-based architecture has reported encouraging results with the global receptive field, the performance and efficiency gaps to the welldeveloped CNN models still exist because of its deteriorative local feature details. This paper proposes a Joint Convolutional Attention and Transformer block (JCAT), which deeply couples the convolutional attention layer and Vision Transformer into one block, as the basic unit to construct our depth completion model in a pyramidal structure. This hybrid architecture naturally benefits both the local connectivity of convolutions and the global context of the Transformer in one single model. As a result, our Completion-Former outperforms state-of-the-art CNNs-based methods on the outdoor KITTI Depth Completion benchmark and indoor NYUv2 dataset, achieving significantly higher efficiency (nearly 1/3 FLOPs) compared to pure Transformerbased methods. Code is available at https://github . com/youmi-zym/CompletionFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers37
- LRRU: Long-short Range Recurrent Updating Networks for Depth CompletionYufei Wang, Bo Li, Ge Zhang, Qi Liu et al.ICCV 2023 · 89 citations
- DepthFM: Fast Generative Monocular Depth Estimation with Flow MatchingMing Gui, Johannes Schusterbauer, Ulrich Prestel, Pingchuan Ma et al.AAAI 2025 · 51 citations
- Depth Anything with Any PriorZehan Wang, Siyu Chen, Lihe Yang, Jialei Wang et al.ICLR 2026 · 47 citations
- Tri-Perspective view Decomposition for Geometry-Aware Depth CompletionZhiqiang Yan, Yuankai Lin, Kun Wang, Yupeng Zheng et al.CVPR 2024 · 33 citations
- Bilateral Propagation Network for Depth CompletionJie Tang, Fei-Peng Tian, Boshi An, Jian Li et al.CVPR 2024 · 33 citations
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
Related papers
- GuideFormer: Transformers for Image Guided Depth CompletionKyeongha Rho, Jinsung Ha, Youngjung KimCVPR 2022 · 57 citations
- DeCoTR: Enhancing Depth Completion with 2D and 3D AttentionsYunxiao Shi, Manish Kumar Singh, Hong Cai, Fatih PorikliCVPR 2024 · 7 citations
- Learning Joint 2D-3D Representations for Depth CompletionYun Chen, Bin Yang, Ming Liang, Raquel UrtasunICCV 2019 · 190 citations
- VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene CompletionYiming Li, Zhiding Yu, Christopher B. Choy, Chaowei Xiao et al.CVPR 2023
- Aggregating Feature Point Cloud for Depth CompletionZhu Yu, Zehua Sheng, Zili Zhou, Lun Luo et al.ICCV 2023 · 42 citations
