Robust Multimodal Depth Estimation using Transformer based Generative Adversarial Networks
Md Fahim Faysal Khan, Anusha Devulapally, Siddharth Advani, Vijaykrishnan Narayanan
摘要
Accurately measuring the absolute depth of every pixel captured by an imaging sensor is of critical importance in real-time applications such as autonomous navigation, augmented reality and robotics. In order to predict dense depth, a general approach is to fuse sensor inputs from different modalities such as LiDAR, camera and other time-of-flight sensors. LiDAR and other time-of-flight sensors provide accurate depth data but are quite sparse, both spatially and temporally. To augment missing depth information, generally RGB guidance is leveraged due to its high resolution information. Due to the reliance on multiple sensor modalities, design for robustness and adaptation is essential. In this work, we propose a transformer-like self-attention based generative adversarial network to estimate dense depth using RGB and sparse depth data. We introduce a novel training recipe for making the model robust so that it works even when one of the input modalities is not available. The multi-head self-attention mechanism can dynamically attend to most salient parts of the RGB image or corresponding sparse depth data producing the most competitive results. Our proposed network also requires less memory for training and inference compared to other existing heavily residual connection based convolutional neural networks, making it more suitable for resource-constrained edge applications. The source code is available at: https://github.com/kocchop/robust-multimodal-fusion-gan
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Sparse to Dense Depth Completion using a Generative Adversarial Network with Intelligent Sampling StrategiesMd Fahim Faysal Khan, Nelson Daniel Troncoso Aldas, Abhishek Kumar, Siddharth Advani 等ACM MM 2021 · 被引用 10 次
- GuideFormer: Transformers for Image Guided Depth CompletionKyeongha Rho, Jinsung Ha, Youngjung KimCVPR 2022 · 被引用 57 次
- RGB-Depth Fusion GAN for Indoor Depth CompletionHaowen Wang, Mingyuan Wang, Zhengping Che, Zhiyuan Xu 等CVPR 2022 · 被引用 47 次
- S3: Learnable Sparse Signal Superdensity for Guided Depth EstimationYu-Kai Huang, Yueh-Cheng Liu, Tsung-Han Wu, Hung-Ting Su 等CVPR 2021
- Sparse Auxiliary Networks for Unified Monocular Depth Prediction and CompletionVitor Guizilini, Rares Ambrus, Wolfram Burgard, Adrien GaidonCVPR 2021
