Exploiting Temporal Consistency for Real-Time Video Depth Estimation
Haokui Zhang, Ying Li, Yuanzhouhan Cao, Yu Liu, Chunhua Shen, Youliang Yan
Abstract
Accuracy of depth estimation from static images has been significantly improved recently, by exploiting hierarchical features from deep convolutional neural networks (CNNs). Compared with static images, vast information exists among video frames and can be exploited to improve the depth estimation performance. In this work, we focus on exploring temporal information from monocular videos for depth estimation. Specifically, we take the advantage of convolutional long short-term memory (CLSTM) and propose a novel spatial-temporal CSLTM (ST-CLSTM) structure. Our ST-CLSTM structure can capture not only the spatial features but also the temporal correlations/consistency among consecutive video frames with negligible increase in computational cost. Additionally, in order to maintain the temporal consistency among the estimated depth frames, we apply the generative adversarial learning scheme and design a temporal consistency loss. The temporal consistency loss is combined with the spatial loss to update the model in an end-to-end fashion. By taking advantage of the temporal information, we build a video depth estimation framework that runs in real-time and generates visually pleasant results. Moreover, our approach is flexible and can be generalized to most existing depth estimation frameworks. Code is available at: https://tinyurl.com/STCLSTM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e772fabf-198c-4a3c-a610-394bfc753436Cited by top-tier papers27
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- Multi-Frame Self-Supervised Depth with TransformersVitor Guizilini, Rares Ambrus, Dian Chen, Sergey Zakharov et al.CVPR 2022 · 95 citations
- MAMo: Leveraging Memory and Attention for Monocular Video Depth EstimationRajeev Yasarla, Hong Cai, Jisoo Jeong, Yunxiao Shi et al.ICCV 2023 · 31 citations
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao et al.ACM MM 2022 · 23 citations
- Instance-wise Occlusion and Depth Orders in Natural ScenesHyunmin Lee, Jaesik ParkCVPR 2022 · 22 citations
Builds on1
Related papers
- Learning Structure Affinity for Video Depth EstimationYuanzhouhan Cao, Yidong Li, Haokui Zhang, Chao Ren et al.ACM MM 2021 · 12 citations
- Self-Supervised Deep Visual Odometry With Online AdaptationShunkai Li, Xin Wang, Yingdian Cao, Fei Xue et al.CVPR 2020
- Spatial Correspondence With Generative Adversarial Network: Learning Depth From Monocular VideosZhenyao Wu, Xinyi Wu, Xiaoping Zhang, Song Wang et al.ICCV 2019 · 29 citations
- Multi-view Depth Estimation using Epipolar Spatio-Temporal NetworksXiaoxiao Long, Lingjie Liu, Wei Li, Christian Theobalt et al.CVPR 2021
- Stereo Any Video: Temporally Consistent Stereo MatchingJunpeng Jing, Weixun Luo, Ye Mao, Krystian MikolajczykICCV 2025 · 1 citation
