Robust Multimodal Depth Estimation using Transformer based Generative Adversarial Networks
Md Fahim Faysal Khan, Anusha Devulapally, Siddharth Advani, Vijaykrishnan Narayanan
Abstract
Accurately measuring the absolute depth of every pixel captured by an imaging sensor is of critical importance in real-time applications such as autonomous navigation, augmented reality and robotics. In order to predict dense depth, a general approach is to fuse sensor inputs from different modalities such as LiDAR, camera and other time-of-flight sensors. LiDAR and other time-of-flight sensors provide accurate depth data but are quite sparse, both spatially and temporally. To augment missing depth information, generally RGB guidance is leveraged due to its high resolution information. Due to the reliance on multiple sensor modalities, design for robustness and adaptation is essential. In this work, we propose a transformer-like self-attention based generative adversarial network to estimate dense depth using RGB and sparse depth data. We introduce a novel training recipe for making the model robust so that it works even when one of the input modalities is not available. The multi-head self-attention mechanism can dynamically attend to most salient parts of the RGB image or corresponding sparse depth data producing the most competitive results. Our proposed network also requires less memory for training and inference compared to other existing heavily residual connection based convolutional neural networks, making it more suitable for resource-constrained edge applications. The source code is available at: https://github.com/kocchop/robust-multimodal-fusion-gan
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 87c596b7-ae4d-43b8-a74a-5d4f317e4a92Related papers
- Sparse to Dense Depth Completion using a Generative Adversarial Network with Intelligent Sampling StrategiesMd Fahim Faysal Khan, Nelson Daniel Troncoso Aldas, Abhishek Kumar, Siddharth Advani et al.ACM MM 2021 · 10 citations
- GuideFormer: Transformers for Image Guided Depth CompletionKyeongha Rho, Jinsung Ha, Youngjung KimCVPR 2022 · 57 citations
- RGB-Depth Fusion GAN for Indoor Depth CompletionHaowen Wang, Mingyuan Wang, Zhengping Che, Zhiyuan Xu et al.CVPR 2022 · 47 citations
- S3: Learnable Sparse Signal Superdensity for Guided Depth EstimationYu-Kai Huang, Yueh-Cheng Liu, Tsung-Han Wu, Hung-Ting Su et al.CVPR 2021
- Sparse Auxiliary Networks for Unified Monocular Depth Prediction and CompletionVitor Guizilini, Rares Ambrus, Wolfram Burgard, Adrien GaidonCVPR 2021
