Exploiting Temporal Consistency for Real-Time Video Depth Estimation
Haokui Zhang, Ying Li, Yuanzhouhan Cao, Yu Liu, Chunhua Shen, Youliang Yan
摘要
Accuracy of depth estimation from static images has been significantly improved recently, by exploiting hierarchical features from deep convolutional neural networks (CNNs). Compared with static images, vast information exists among video frames and can be exploited to improve the depth estimation performance. In this work, we focus on exploring temporal information from monocular videos for depth estimation. Specifically, we take the advantage of convolutional long short-term memory (CLSTM) and propose a novel spatial-temporal CSLTM (ST-CLSTM) structure. Our ST-CLSTM structure can capture not only the spatial features but also the temporal correlations/consistency among consecutive video frames with negligible increase in computational cost. Additionally, in order to maintain the temporal consistency among the estimated depth frames, we apply the generative adversarial learning scheme and design a temporal consistency loss. The temporal consistency loss is combined with the spatial loss to update the model in an end-to-end fashion. By taking advantage of the temporal information, we build a video depth estimation framework that runs in real-time and generates visually pleasant results. Moreover, our approach is flexible and can be generalized to most existing depth estimation frameworks. Code is available at: https://tinyurl.com/STCLSTM
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
- Multi-Frame Self-Supervised Depth with TransformersVitor Guizilini, Rares Ambrus, Dian Chen, Sergey Zakharov 等CVPR 2022 · 被引用 95 次
- MAMo: Leveraging Memory and Attention for Monocular Video Depth EstimationRajeev Yasarla, Hong Cai, Jisoo Jeong, Yunxiao Shi 等ICCV 2023 · 被引用 31 次
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao 等ACM MM 2022 · 被引用 23 次
- Instance-wise Occlusion and Depth Orders in Natural ScenesHyunmin Lee, Jaesik ParkCVPR 2022 · 被引用 22 次
它引用的顶会 Paper1
相关 Paper
- Learning Structure Affinity for Video Depth EstimationYuanzhouhan Cao, Yidong Li, Haokui Zhang, Chao Ren 等ACM MM 2021 · 被引用 12 次
- Self-Supervised Deep Visual Odometry With Online AdaptationShunkai Li, Xin Wang, Yingdian Cao, Fei Xue 等CVPR 2020
- Spatial Correspondence With Generative Adversarial Network: Learning Depth From Monocular VideosZhenyao Wu, Xinyi Wu, Xiaoping Zhang, Song Wang 等ICCV 2019 · 被引用 29 次
- Multi-view Depth Estimation using Epipolar Spatio-Temporal NetworksXiaoxiao Long, Lingjie Liu, Wei Li, Christian Theobalt 等CVPR 2021
- Stereo Any Video: Temporally Consistent Stereo MatchingJunpeng Jing, Weixun Luo, Ye Mao, Krystian MikolajczykICCV 2025 · 被引用 1 次
