MVS2D: Efficient Multiview Stereo via Attention-Driven 2D Convolutions
Zhenpei Yang, Zhile Ren, Qi Shan, Qixing Huang
Abstract
Deep learning has made significant impacts on multiview stereo systems. State-of-the-art approaches typically involve building a cost volume, followed by multiple 3D convolution operations to recover the input image's pixel-wise depth. While such end-to-end learning of plane-sweeping stereo advances public benchmarks' accuracy, they are typically very slow to compute. We present MVS2D, a highly efficient multi-view stereo algorithm that seamlessly integrates multi-view constraints into single-view net-works via an attention mechanism. Since MVS2D only builds on 2D convolutions, it is at least <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> than all the notable counterparts. Moreover, our algorithm produces precise depth estimations and 3D reconstructions, achieving state-of-the-art results on challenging benchmarks ScanNet, SUN3D, RGBD, and the classical DTU dataset. our algorithm also outperforms all other algorithms in the setting of inexact camera poses. Our code is released at https://github.com/zhenpeiyang/MVS2D
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c6cc472-80ce-43af-adca-ca1282bdfaa0Cited by top-tier papers19
- WT-MVSNet: Window-based Transformers for Multi-view StereoJinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang et al.NeurIPS 2022 · 50 citations
- Input-level Inductive Biases for 3D ReconstructionWang Yifan, Carl Doersch, Relja Arandjelovic, João Carreira et al.CVPR 2022 · 23 citations
- Rethinking Disparity: A Depth Range Free Multi-View Stereo Based on DisparityQingsong Yan, Qiang Wang, Kaiyong Zhao, Bo Li et al.AAAI 2023 · 21 citations
- Test3R: Learning to Reconstruct 3D at Test TimeYuheng Yuan, Qiuhong Shen, Shizun Wang, Xingyi Yang et al.NeurIPS 2025 · 19 citations
- Is Attention All That NeRF Needs?Mukund Varma T., Peihao Wang, Xuxi Chen, Tianlong Chen et al.ICLR 2023 · 6 citations
Builds on16
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 487 citations
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 403 citations
- Learning Inverse Depth Regression for Multi-View Stereo with Correlation Cost VolumeQingshan Xu, Wenbing TaoAAAI 2020 · 145 citations
- Cost Volume Pyramid Based Depth Inference for Multi-View StereoJiayu Yang, Wei Mao, José M. Álvarez, Miaomiao LiuCVPR 2020
- StruMonoNet: Structure-Aware Monocular 3D PredictionZhenpei Yang, Li Erran Li, Qixing HuangCVPR 2021
Related papers
- P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Haipeng Huang et al.ICCV 2019 · 254 citations
- A Confidence-based Iterative Solver of Depths and Surface Normals for Deep Multi-view StereoWang Zhao, Shaohui Liu, Yi Wei, Hengkai Guo et al.ICCV 2021 · 16 citations
- Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton RefinementZehao Yu, Shenghua GaoCVPR 2020
- VolumeFusion: Deep Depth Fusion for 3D Scene ReconstructionJaesung Choe, Sunghoon Im, François Rameau, Minjun Kang et al.ICCV 2021 · 83 citations
- MVSCRF: Learning Multi-View Stereo With Conditional Random FieldsYouze Xue, Jiansheng Chen, Weitao Wan, Yiqing Huang et al.ICCV 2019 · 95 citations
