CompletionFormer: Depth Completion with Convolutions and Vision Transformers
Youmin Zhang, Xianda Guo, Matteo Poggi, Zheng Zhu, Guan Huang, Stefano Mattoccia
摘要
Given sparse depths and the corresponding RGB images, depth completion aims at spatially propagating the sparse measurements throughout the whole image to get a dense depth prediction. Despite the tremendous progress of deeplearning-based depth completion methods, the locality of the convolutional layer or graph model makes it hard for the network to model the long-range relationship between pixels. While recent fully Transformer-based architecture has reported encouraging results with the global receptive field, the performance and efficiency gaps to the welldeveloped CNN models still exist because of its deteriorative local feature details. This paper proposes a Joint Convolutional Attention and Transformer block (JCAT), which deeply couples the convolutional attention layer and Vision Transformer into one block, as the basic unit to construct our depth completion model in a pyramidal structure. This hybrid architecture naturally benefits both the local connectivity of convolutions and the global context of the Transformer in one single model. As a result, our Completion-Former outperforms state-of-the-art CNNs-based methods on the outdoor KITTI Depth Completion benchmark and indoor NYUv2 dataset, achieving significantly higher efficiency (nearly 1/3 FLOPs) compared to pure Transformerbased methods. Code is available at https://github . com/youmi-zym/CompletionFormer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- LRRU: Long-short Range Recurrent Updating Networks for Depth CompletionYufei Wang, Bo Li, Ge Zhang, Qi Liu 等ICCV 2023 · 被引用 89 次
- DepthFM: Fast Generative Monocular Depth Estimation with Flow MatchingMing Gui, Johannes Schusterbauer, Ulrich Prestel, Pingchuan Ma 等AAAI 2025 · 被引用 51 次
- Depth Anything with Any PriorZehan Wang, Siyu Chen, Lihe Yang, Jialei Wang 等ICLR 2026 · 被引用 47 次
- Tri-Perspective view Decomposition for Geometry-Aware Depth CompletionZhiqiang Yan, Yuankai Lin, Kun Wang, Yupeng Zheng 等CVPR 2024 · 被引用 33 次
- Bilateral Propagation Network for Depth CompletionJie Tang, Fei-Peng Tian, Boshi An, Jian Li 等CVPR 2024 · 被引用 33 次
它引用的顶会 Paper20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
相关 Paper
- GuideFormer: Transformers for Image Guided Depth CompletionKyeongha Rho, Jinsung Ha, Youngjung KimCVPR 2022 · 被引用 57 次
- DeCoTR: Enhancing Depth Completion with 2D and 3D AttentionsYunxiao Shi, Manish Kumar Singh, Hong Cai, Fatih PorikliCVPR 2024 · 被引用 7 次
- Learning Joint 2D-3D Representations for Depth CompletionYun Chen, Bin Yang, Ming Liang, Raquel UrtasunICCV 2019 · 被引用 190 次
- VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene CompletionYiming Li, Zhiding Yu, Christopher B. Choy, Chaowei Xiao 等CVPR 2023
- Aggregating Feature Point Cloud for Depth CompletionZhu Yu, Zehua Sheng, Zili Zhou, Lun Luo 等ICCV 2023 · 被引用 42 次
