GuideFormer: Transformers for Image Guided Depth Completion
Kyeongha Rho, Jinsung Ha, Youngjung Kim
摘要
Depth completion has been widely studied to predict a dense depth image from its sparse measurement and a single color image. However, most state-of-the-art methods rely on static convolutional neural networks (CNNs) which are not flexible enough for capturing the dynamic nature of input contexts. In this paper, we propose GuideFormer, a fully transformer-based architecture for dense depth completion. We first process sparse depth and color guidance images with separate transformer branches to extract hierarchical and complementary token representations. Each branch consists of a stack of self-attention blocks and has key design features to make our model suitable for the task. We also devise an effective token fusion method based on guided-attention mechanism. It explicitly models information flow between the two branches and captures inter-modal dependencies that cannot be obtained from depth or color image alone. These properties allow GuideFormer to enjoy various visual dependencies and recover precise depth values while preserving fine details. We evaluate GuideFormer on the KITTI dataset containing realworld driving scenes and provide extensive ablation studies. Experimental results demonstrate that our approach significantly outperforms the state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- LRRU: Long-short Range Recurrent Updating Networks for Depth CompletionYufei Wang, Bo Li, Ge Zhang, Qi Liu 等ICCV 2023 · 被引用 89 次
- DesNet: Decomposed Scale-Consistent Network for Unsupervised Depth CompletionZhiqiang Yan, Kun Wang, Xiang Li, Zhenyu Zhang 等AAAI 2023 · 被引用 46 次
- Aggregating Feature Point Cloud for Depth CompletionZhu Yu, Zehua Sheng, Zili Zhou, Lun Luo 等ICCV 2023 · 被引用 42 次
- Tri-Perspective view Decomposition for Geometry-Aware Depth CompletionZhiqiang Yan, Yuankai Lin, Kun Wang, Yupeng Zheng 等CVPR 2024 · 被引用 33 次
- Bilateral Propagation Network for Depth CompletionJie Tang, Fei-Peng Tian, Boshi An, Jian Li 等CVPR 2024 · 被引用 33 次
它引用的顶会 Paper13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
相关 Paper
- CompletionFormer: Depth Completion with Convolutions and Vision TransformersYoumin Zhang, Xianda Guo, Matteo Poggi, Zheng Zhu 等CVPR 2023
- VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene CompletionYiming Li, Zhiding Yu, Christopher B. Choy, Chaowei Xiao 等CVPR 2023
- FCFR-Net: Feature Fusion based Coarse-to-Fine Residual Learning for Depth CompletionLina Liu, Xibin Song, Xiaoyang Lyu, Junwei Diao 等AAAI 2021 · 被引用 125 次
- Robust Multimodal Depth Estimation using Transformer based Generative Adversarial NetworksMd Fahim Faysal Khan, Anusha Devulapally, Siddharth Advani, Vijaykrishnan NarayananACM MM 2022 · 被引用 6 次
- Multi-Frame Self-Supervised Depth with TransformersVitor Guizilini, Rares Ambrus, Dian Chen, Sergey Zakharov 等CVPR 2022 · 被引用 95 次
