Generalized Binary Search Network for Highly-Efficient Multi-View Stereo
Zhenxing Mi, Di Chang, Dan Xu
摘要
Multi-view Stereo (MVS) with known camera parameters is essentially a 1D search problem within a valid depth range. Recent deep learning-based MVS methods typically densely sample depth hypotheses in the depth range, and then construct prohibitively memory-consuming 3D cost volumes for depth prediction. Although coarse-to-fine sampling strategies alleviate this overhead issue to a certain extent, the efficiency of MVS is still an open challenge. In this work, we propose a novel method for highly efficient MVS that remarkably decreases the memory footprint, meanwhile clearly advancing state-of-the-art depth prediction performance. We investigate what a search strategy can be reasonably optimal for MVS taking into account of both efficiency and effectiveness. We first formulate MVS as a binary search problem, and accordingly propose a generalized binary search network for MVS. Specifically, in each step, the depth range is split into 2 bins with extra 1 error tolerance bin on both sides. A classification is performed to identify which bin contains the true depth. We also design three mechanisms to respectively handle classification errors, deal with out-of-range samples and decrease the training memory. The new formulation makes our method only sample a very small number of depth hypotheses in each step, which is highly memory efficient, and also greatly facilitates quick training convergence. Experiments on competitive benchmarks show that our method achieves state-of-the-art accuracy with much less memory. Particularly, our method obtains an overall score of 0.289 on DTU dataset and tops the first place on challenging Tanks and Temples advanced dataset among all the learning-based methods. Our code will be released at https://github.com/MiZhenxing/GBi-Net .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View StereoChenjie Cao, Xinlin Ren, Yanwei FuICLR 2024 · 被引用 68 次
- WT-MVSNet: Window-based Transformers for Multi-view StereoJinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang 等NeurIPS 2022 · 被引用 50 次
- When Epipolar Constraint Meets Non-local Operators in Multi-View StereoTianqi Liu, Xinyi Ye, Weiyue Zhao, Zhiyu Pan 等ICCV 2023 · 被引用 43 次
- Efficient Edge-Preserving Multi-View Stereo Network for Depth EstimationWanjuan Su, Wenbing TaoAAAI 2023 · 被引用 40 次
- GoMVS: Geometrically Consistent Cost Aggregation for Multi-View StereoJiang Wu, Rui Li, Haofei Xu, Wenxun Zhao 等CVPR 2024 · 被引用 34 次
它引用的顶会 Paper11
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 被引用 403 次
- AA-RMVSNet: Adaptive Aggregation Recurrent Multi-view Stereo NetworkZizhuang Wei, Qingtian Zhu, Chen Min, Yisong Chen 等ICCV 2021 · 被引用 193 次
- EPP-MVSNet: Epipolar-assembling based Depth Prediction for Multi-view StereoXinjun Ma, Yue Gong, Qirui Wang, Jingwei Huang 等ICCV 2021 · 被引用 147 次
- Learning Inverse Depth Regression for Multi-View Stereo with Correlation Cost VolumeQingshan Xu, Wenbing TaoAAAI 2020 · 被引用 145 次
- MVS2D: Efficient Multiview Stereo via Attention-Driven 2D ConvolutionsZhenpei Yang, Zhile Ren, Qi Shan, Qixing HuangCVPR 2022 · 被引用 43 次
相关 Paper
- Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton RefinementZehao Yu, Shenghua GaoCVPR 2020
- Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo MatchingXiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai 等CVPR 2020
- RayMVSNet: Learning Ray-based 1D Implicit Fields for Accurate Multi-View StereoJunhua Xi, Yifei Shi, Yijie Wang, Yulan Guo 等CVPR 2022 · 被引用 129 次
- Non-parametric Depth Distribution Modelling based Depth Inference for Multi-view StereoJiayu Yang, José M. Álvarez, Miaomiao LiuCVPR 2022 · 被引用 39 次
- Constraining Depth Map Geometry for Multi-View Stereo: A Dual-Depth Approach with Saddle-shaped Depth CellsXinyi Ye, Weiyue Zhao, Tianqi Liu, Zihao Huang 等ICCV 2023 · 被引用 29 次
