Multi-view Depth Estimation using Epipolar Spatio-Temporal Networks
Xiaoxiao Long, Lingjie Liu, Wei Li, Christian Theobalt, Wenping Wang
摘要
We present a novel method for multi-view depth estimation from a single video, which is a critical task in various applications, such as perception, reconstruction and robot navigation. Although previous learning-based methods have demonstrated compelling results, most works estimate depth maps of individual video frames independently, without taking into consideration the strong geometric and temporal coherence among the frames. Moreover, current state-of-the-art (SOTA) models mostly adopt a fully 3D convolution network for cost regularization and therefore require high computational cost, thus limiting their deployment in real-world applications. Our method achieves temporally coherent depth estimation results by using a novel Epipolar Spatio-Temporal (EST) transformer to explicitly associate geometric and temporal correlation with multiple estimated depth maps. Furthermore, to reduce the computational cost, inspired by recent Mixture-of-Experts models, we design a compact hybrid network consisting of a 2D context-aware network and a 3D matching network which learn 2D context information and 3D disparity cues separately. Extensive experiments demonstrate that our method achieves higher accuracy in depth estimation and significant speedup than the SOTA methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- GaussianPro: 3D Gaussian Splatting with Progressive PropagationKai Cheng, Xiaoxiao Long, Kaizhi Yang, Yao Yao 等ICML 2024 · 被引用 241 次
- Multi-View Depth Estimation by Fusing Single-View Depth Probability with Multi-View GeometryGwangbin Bae, Ignas Budvytis, Roberto CipollaCVPR 2022 · 被引用 58 次
- PlaneMVS: 3D Plane Reconstruction from Multi-View StereoJiachen Liu, Pan Ji, Nitin Bansal, Changjiang Cai 等CVPR 2022 · 被引用 43 次
- MVS2D: Efficient Multiview Stereo via Attention-Driven 2D ConvolutionsZhenpei Yang, Zhile Ren, Qi Shan, Qixing HuangCVPR 2022 · 被引用 43 次
- PlanarRecon: Realtime 3D Plane Detection and Reconstruction from Posed Monocular VideosYiming Xie, Matheus Gadelha, Fengting Yang, Xiaowei Zhou 等CVPR 2022 · 被引用 32 次
它引用的顶会 Paper6
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 被引用 487 次
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
- Depth Completion From Sparse LiDAR Data With Depth-Normal ConstraintsYan Xu, Xinge Zhu, Jianping Shi, Guofeng Zhang 等ICCV 2019 · 被引用 249 次
- Learning a Mixture of Granularity-Specific Experts for Fine-Grained CategorizationLianbo Zhang, Shaoli Huang, Wei Liu, Dacheng TaoICCV 2019 · 被引用 191 次
- Exploiting Temporal Consistency for Real-Time Video Depth EstimationHaokui Zhang, Ying Li, Yuanzhouhan Cao, Yu Liu 等ICCV 2019 · 被引用 137 次
相关 Paper
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao 等ACM MM 2022 · 被引用 23 次
- Multi-Frame Self-Supervised Depth with TransformersVitor Guizilini, Rares Ambrus, Dian Chen, Sergey Zakharov 等CVPR 2022 · 被引用 95 次
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang 等CVPR 2022 · 被引用 403 次
- Deep Semantic Graph Transformer for Multi-View 3D Human Pose EstimationLijun Zhang, Kangkang Zhou, Feng Lu, Xiang-Dong Zhou 等AAAI 2024 · 被引用 14 次
- WT-MVSNet: Window-based Transformers for Multi-view StereoJinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang 等NeurIPS 2022 · 被引用 50 次
