Learning Structure Affinity for Video Depth Estimation
Yuanzhouhan Cao, Yidong Li, Haokui Zhang, Chao Ren, Yifan Liu
Abstract
Depth estimation is a structure learning problem. The affinity among neighbouring pixels plays an important role in inferring depth values. In this paper, we propose to learn structure affinity in both spatial and temporal domain for accurate depth estimation from monocular videos. Specifically, we first propose a convolutional spatial temporal propagation network (CSTPN) that learns affinity among neighbouring video frames. Secondly, we employ a structure knowledge distillation scheme that transfers the spatial temporal affinity learned by cumbersome network to compact network. By calculating pixel-wise similarities between neighboring frames and neighbouring sequences, our knowledge distillation scheme efficiently captures both short-term and long-term spatial temporal affinity. Finally, we apply a warping loss based on optical flow between video frames to further enforce the temporal affinity. Experiment results show that our proposed depth estimation approach outperform the state-of-the-art methods on both indoor and outdoor benchmark datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 085549d3-409f-42a6-8fc5-021a4eb791d6Cited by top-tier papers8
- MAMo: Leveraging Memory and Attention for Monocular Video Depth EstimationRajeev Yasarla, Hong Cai, Jisoo Jeong, Yunxiao Shi et al.ICCV 2023 · 31 citations
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao et al.ACM MM 2022 · 23 citations
- Digging into Depth Priors for Outdoor Neural Radiance FieldsChen Wang, Jiadai Sun, Lina Liu, Chenming Wu et al.ACM MM 2023 · 13 citations
- Learning Temporally Consistent Video Depth from Video Diffusion PriorsJiahao Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang et al.CVPR 2025
- Neural Video Depth StabilizerYiran Wang, Min Shi, Jiaqi Li, Zihao Huang et al.ICCV 2023
Related papers
- Exploiting Temporal Consistency for Real-Time Video Depth EstimationHaokui Zhang, Ying Li, Yuanzhouhan Cao, Yu Liu et al.ICCV 2019 · 137 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- Temporally Distributed Networks for Fast Video Semantic SegmentationPing Hu, Fabian Caba, Oliver Wang, Zhe Lin et al.CVPR 2020
- Spatial Correspondence With Generative Adversarial Network: Learning Depth From Monocular VideosZhenyao Wu, Xinyi Wu, Xiaoping Zhang, Song Wang et al.ICCV 2019 · 29 citations
- Stereo Any Video: Temporally Consistent Stereo MatchingJunpeng Jing, Weixun Luo, Ye Mao, Krystian MikolajczykICCV 2025 · 1 citation
