Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton Refinement
Zehao Yu, Shenghua Gao
Abstract
Almost all previous deep learning-based multi-view stereo (MVS) approaches focus on improving reconstruction quality. Besides quality, efficiency is also a desirable feature for MVS in real scenarios. Towards this end, this paper presents a Fast-MVSNet, a novel sparse-to-dense coarse-to-fine framework, for fast and accurate depth estimation in MVS. Specifically, in our Fast-MVSNet, we first construct a sparse cost volume for learning a sparse and high-resolution depth map. Then we leverage a smallscale convolutional neural network to encode the depth dependencies for pixels within a local region to densify the sparse high-resolution depth map. At last, a simple but efficient Gauss-Newton layer is proposed to further optimize the depth map. On one hand, the high-resolution depth map, the data-adaptive propagation method and the Gauss-Newton layer jointly guarantee the effectiveness of our method. On the other hand, all modules in our Fast-MVSNet are lightweight and thus guarantee the efficiency of our approach. Besides, our approach is also memoryfriendly because of the sparse depth representation. Extensive experimental results show that our method is 5× and 14× faster than Point-MVSNet and R-MVSNet, respectively, while achieving comparable or even better results on the challenging Tanks and Temples dataset as well as the DTU dataset. Code is available at https://github.com/ svip-lab/FastMVSNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85be6c21-e6fa-4349-b6e0-c185e279b62eCited by top-tier papers47
- MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface ReconstructionZehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler et al.NeurIPS 2022 · 670 citations
- 2D Gaussian Splatting for Geometrically Accurate Radiance FieldsBinbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger et al.SIGGRAPH 2024 · 660 citations
- Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion PriorsGuocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren et al.ICLR 2024 · 444 citations
- Pixel-Perfect Structure-from-Motion with Featuremetric RefinementPhilipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, Marc PollefeysICCV 2021 · 266 citations
- BEVStereo: Enhancing Depth Estimation in Multi-View 3D Object Detection with Temporal StereoYinhao Li, Han Bao, Zheng Ge, Jinrong Yang et al.AAAI 2023 · 226 citations
Builds on1
Related papers
- Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo MatchingXiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai et al.CVPR 2020
- Generalized Binary Search Network for Highly-Efficient Multi-View StereoZhenxing Mi, Di Chang, Dan XuCVPR 2022
- Efficient Multi-view Stereo by Iterative Dynamic Cost VolumeShaoqian Wang, Bo Li, Yuchao DaiCVPR 2022 · 62 citations
- RayMVSNet: Learning Ray-based 1D Implicit Fields for Accurate Multi-View StereoJunhua Xi, Yifei Shi, Yijie Wang, Yulan Guo et al.CVPR 2022 · 129 citations
- P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Haipeng Huang et al.ICCV 2019 · 254 citations
