Learning Structure Affinity for Video Depth Estimation
Yuanzhouhan Cao, Yidong Li, Haokui Zhang, Chao Ren, Yifan Liu
摘要
Depth estimation is a structure learning problem. The affinity among neighbouring pixels plays an important role in inferring depth values. In this paper, we propose to learn structure affinity in both spatial and temporal domain for accurate depth estimation from monocular videos. Specifically, we first propose a convolutional spatial temporal propagation network (CSTPN) that learns affinity among neighbouring video frames. Secondly, we employ a structure knowledge distillation scheme that transfers the spatial temporal affinity learned by cumbersome network to compact network. By calculating pixel-wise similarities between neighboring frames and neighbouring sequences, our knowledge distillation scheme efficiently captures both short-term and long-term spatial temporal affinity. Finally, we apply a warping loss based on optical flow between video frames to further enforce the temporal affinity. Experiment results show that our proposed depth estimation approach outperform the state-of-the-art methods on both indoor and outdoor benchmark datasets.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- MAMo: Leveraging Memory and Attention for Monocular Video Depth EstimationRajeev Yasarla, Hong Cai, Jisoo Jeong, Yunxiao Shi 等ICCV 2023 · 被引用 31 次
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao 等ACM MM 2022 · 被引用 23 次
- Digging into Depth Priors for Outdoor Neural Radiance FieldsChen Wang, Jiadai Sun, Lina Liu, Chenming Wu 等ACM MM 2023 · 被引用 13 次
- Learning Temporally Consistent Video Depth from Video Diffusion PriorsJiahao Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang 等CVPR 2025
- Neural Video Depth StabilizerYiran Wang, Min Shi, Jiaqi Li, Zihao Huang 等ICCV 2023
相关 Paper
- Exploiting Temporal Consistency for Real-Time Video Depth EstimationHaokui Zhang, Ying Li, Yuanzhouhan Cao, Yu Liu 等ICCV 2019 · 被引用 137 次
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
- Temporally Distributed Networks for Fast Video Semantic SegmentationPing Hu, Fabian Caba, Oliver Wang, Zhe Lin 等CVPR 2020
- Spatial Correspondence With Generative Adversarial Network: Learning Depth From Monocular VideosZhenyao Wu, Xinyi Wu, Xiaoping Zhang, Song Wang 等ICCV 2019 · 被引用 29 次
- Stereo Any Video: Temporally Consistent Stereo MatchingJunpeng Jing, Weixun Luo, Ye Mao, Krystian MikolajczykICCV 2025 · 被引用 1 次
