Spatial Correspondence With Generative Adversarial Network: Learning Depth From Monocular Videos
Zhenyao Wu, Xinyi Wu, Xiaoping Zhang, Song Wang, Lili Ju
Abstract
Depth estimation from monocular videos has important applications in many areas such as autonomous driving and robot navigation. It is a very challenging problem without knowing the camera pose since errors in camera-pose estimation can significantly affect the video-based depth estimation accuracy. In this paper, we present a novel SC-GAN network with end-to-end adversarial training for depth estimation from monocular videos without estimating the camera pose and pose change over time. To exploit cross-frame relations, SC-GAN includes a spatial correspondence module which uses Smolyak sparse grids to efficiently match the features across adjacent frames, and an attention mechanism to learn the importance of features in different directions. Furthermore, the generator in SC-GAN learns to estimate depth from the input frames, while the discriminator learns to distinguish between the ground-truth and estimated depth map for the reference frame. Experiments on the KITTI and Cityscapes datasets show that the proposed SC-GAN can achieve much more accurate depth maps than many existing state-of-the-art methods on monocular videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42a96fa5-7c0c-45fa-b3a5-b9f6360c4a50Cited by top-tier papers7
- Multi-Frame Self-Supervised Depth with TransformersVitor Guizilini, Rares Ambrus, Dian Chen, Sergey Zakharov et al.CVPR 2022 · 95 citations
- Multi-View Depth Estimation by Fusing Single-View Depth Probability with Multi-View GeometryGwangbin Bae, Ignas Budvytis, Roberto CipollaCVPR 2022 · 58 citations
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao et al.ACM MM 2022 · 23 citations
- Multi-Resolution Monocular Depth Map Fusion by Self-Supervised Gradient-Based CompositionYaqiao Dai, Renjiao Yi, Chenyang Zhu, Hongjun He et al.AAAI 2023 · 8 citations
- Adaptive Fusion of Single-View and Multi-View Depth for Autonomous DrivingJunda Cheng, Wei Yin, Kaixuan Wang, Xiaozhi Chen et al.CVPR 2024
Builds on1
Related papers
- Sequential Adversarial Learning for Self-Supervised Deep Visual OdometryShunkai Li, Fei Xue, Xin Wang, Zike Yan et al.ICCV 2019 · 58 citations
- Learning Structure Affinity for Video Depth EstimationYuanzhouhan Cao, Yidong Li, Haokui Zhang, Chao Ren et al.ACM MM 2021 · 12 citations
- Patch-Wise Attention Network for Monocular Depth EstimationSihaeng Lee, Janghyeon Lee, Byungju Kim, Eojindl Yi et al.AAAI 2021 · 84 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- Exploiting Temporal Consistency for Real-Time Video Depth EstimationHaokui Zhang, Ying Li, Yuanzhouhan Cao, Yu Liu et al.ICCV 2019 · 137 citations
