Non-parametric Depth Distribution Modelling based Depth Inference for Multi-view Stereo
Jiayu Yang, José M. Álvarez, Miaomiao Liu
Abstract
Recent cost volume pyramid based deep neural networks have unlocked the potential of efficiently leveraging high-resolution images for depth inference from multi-view stereo. In general, those approaches assume that the depth of each pixel follows a unimodal distribution. Boundary pixels usually follow a multi-modal distribution as they represent different depths; Therefore, the assumption results in an erroneous depth prediction at the coarser level of the cost volume pyramid and can not be corrected in the refinement levels leading to wrong depth predictions. In contrast, we propose constructing the cost volume by non-parametric depth distribution modeling to handle pixels with unimodal and multi-modal distributions. Our approach outputs multiple depth hypotheses at the coarser level to avoid errors in the early stage. As we perform local search around these multiple hypotheses in subsequent levels, our approach does not maintain the rigid depth spatial ordering and, therefore, we introduce a sparse cost aggregation network to derive information within each volume. We evaluate our approach extensively on two benchmark datasets: DTU and Tanks & Temples. Our experimental results show that our model outperforms existing methods by a large margin and achieves superior performance on boundary regions. Code is available at https://github.com/NVlabs/NP-CVP-MVSNet
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View StereoChenjie Cao, Xinlin Ren, Yanwei FuICLR 2024 · 68 citations
- Efficient Edge-Preserving Multi-View Stereo Network for Depth EstimationWanjuan Su, Wenbing TaoAAAI 2023 · 40 citations
- GoMVS: Geometrically Consistent Cost Aggregation for Multi-View StereoJiang Wu, Rui Li, Haofei Xu, Wenxun Zhao et al.CVPR 2024 · 34 citations
- Adaptive Multi-Modal Cross-Entropy Loss for Stereo MatchingPeng Xu, Zhiyu Xiang, Chengyu Qiao, Jingyun Fu et al.CVPR 2024 · 28 citations
- CL-MVSNet: Unsupervised Multi-view Stereo with Dual-level Contrastive LearningKaiqiang Xiong, Rui Peng, Zhe Zhang, Tianxing Feng et al.ICCV 2023 · 24 citations
Builds on11
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 403 citations
- AA-RMVSNet: Adaptive Aggregation Recurrent Multi-view Stereo NetworkZizhuang Wei, Qingtian Zhu, Chen Min, Yisong Chen et al.ICCV 2021 · 193 citations
- EPP-MVSNet: Epipolar-assembling based Depth Prediction for Multi-view StereoXinjun Ma, Yue Gong, Qirui Wang, Jingwei Huang et al.ICCV 2021 · 147 citations
- PatchMatch-RL: Deep MVS with Pixelwise Depth, Normal, and VisibilityJae Yong Lee, Joseph DeGol, Chuhang Zou, Derek HoiemICCV 2021 · 35 citations
- UASNet: Uncertainty Adaptive Sampling Network for Deep Stereo MatchingYamin Mao, Zhihua Liu, Weiming Li, Yuchao Dai et al.ICCV 2021 · 34 citations
Related papers
- Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo MatchingXiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai et al.CVPR 2020
- P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Haipeng Huang et al.ICCV 2019 · 254 citations
- Multi-View Stereo Representation Revist: Region-Aware MVSNetYisu Zhang, Jianke Zhu, Lixiang LinCVPR 2023
- Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton RefinementZehao Yu, Shenghua GaoCVPR 2020
- Generalized Binary Search Network for Highly-Efficient Multi-View StereoZhenxing Mi, Di Chang, Dan XuCVPR 2022
