Just a Few Points are All You Need for Multi-view Stereo: A Novel Semi-supervised Learning Method for Multi-view Stereo
Taekyung Kim, Jaehoon Choi, Seokeon Choi, Dongki Jung, Changick Kim
Abstract
While learning-based multi-view stereo (MVS) methods have recently shown successful performances in quality and efficiency, limited MVS data hampers generalization to unseen environments. A simple solution is to generate various large-scale MVS datasets, but generating dense ground truth for 3D structure requires a huge amount of time and resources. On the other hand, if the reliance on dense ground truth is relaxed, MVS systems will generalize more smoothly to new environments. To this end, we first introduce a novel semi-supervised multi-view stereo framework called a Sparse Ground truth-based MVS Network (SGT-MVSNet) that can reliably reconstruct the 3D structures even with a few ground truth 3D points. Our strategy is to divide the accurate and erroneous regions and individually conquer them based on our observation that a probability map can separate these regions. We propose a self-supervision loss called the 3D Point Consistency Loss to enhance the 3D reconstruction performance, which forces the 3D points back-projected from the corresponding pixels by the predicted depth values to meet at the same 3D co-ordinates. Finally, we propagate these improved depth pre-dictions toward edges and occlusions by the Coarse-to-fine Reliable Depth Propagation module. We generate the spare ground truth of the DTU dataset for evaluation and extensive experiments verify that our SGT-MVSNet outperforms the state-of-the-art MVS methods on the sparse ground truth setting. Moreover, our method shows comparable reconstruction results to the supervised MVS methods though we only used tens and hundreds of ground truth 3D points.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccbc63e2-8b91-4fc1-a81c-21a4ebed7582Cited by top-tier papers2
- Semi-supervised Deep Multi-view StereoHongbin Xu, Weitao Chen, Yang Liu, Zhipeng Zhou et al.ACM MM 2023 · 7 citations
- UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery Using Gaussian SplattingJaehoon Choi, Dongki Jung, Chris Maxey, Sungmin Eum et al.AAAI 2026 · 2 citations
Builds on3
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 403 citations
- Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton RefinementZehao Yu, Shenghua GaoCVPR 2020
- Attention-Aware Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Yuesong Wang et al.CVPR 2020
Related papers
- Self-supervised Multi-view Stereo via Inter and Intra Network Pseudo DepthKe Qiu, Yawen Lai, Shiyi Liu, Ronggang WangACM MM 2022 · 9 citations
- Self-supervised Multi-view Stereo via Effective Co-Segmentation and Data-AugmentationHongbin Xu, Zhipeng Zhou, Yu Qiao, Wenxiong Kang et al.AAAI 2021 · 86 citations
- CL-MVSNet: Unsupervised Multi-view Stereo with Dual-level Contrastive LearningKaiqiang Xiong, Rui Peng, Zhe Zhang, Tianxing Feng et al.ICCV 2023 · 24 citations
- Digging into Uncertainty in Self-supervised Multi-view StereoHongbin Xu, Zhipeng Zhou, Yali Wang, Wenxiong Kang et al.ICCV 2021 · 68 citations
- DS-MVSNet: Unsupervised Multi-view Stereo via Depth SynthesisJingliang Li, Zhengda Lu, Yiqun Wang, Ying Wang et al.ACM MM 2022 · 19 citations
