A Novel Recurrent Encoder-Decoder Structure for Large-Scale Multi-View Stereo Reconstruction From an Open Aerial Dataset
Jin Liu, Shunping Ji
Abstract
A great deal of research has demonstrated recently that multi-view stereo (MVS) matching can be solved with deep learning methods. However, these efforts were focused on close-range objects and only a very few of the deep learning-based methods were specifically designed for large-scale 3D urban reconstruction due to the lack of multi-view aerial image benchmarks. In this paper, we present a synthetic aerial dataset, called the WHU dataset, we created for MVS tasks, which, to our knowledge, is the first large-scale multi-view aerial dataset. It was generated from a highly accurate 3D digital surface model produced from thousands of real aerial images with precise camera parameters. We also introduce in this paper a novel network, called RED-Net, for wide-range depth inference, which we developed from a recurrent encoderdecoder structure to regularize cost maps across depths and a 2D fully convolutional network as framework. RED-Net's low memory requirements and high performance make it suitable for large-scale and highly accurate 3D Earth surface reconstruction. Our experiments confirmed that not only did our method exceed the current state-of-the-art MVS methods by more than 50% mean absolute error (MAE) with less memory and computational cost, but its efficiency as well. It outperformed one of the best commercial software programs based on conventional methods, improving their efficiency 16 times over. Moreover, we proved that our RED-Net model pre-trained on the synthetic WHU dataset can be efficiently transferred to very different multi-view aerial image datasets without any fine-tuning. Dataset and code are available at http://gpcv.whu.edu.cn/data .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68ffcc67-9446-4dde-b707-8ba82e1c4f08Cited by top-tier papers8
- Rational Polynomial Camera Model Warping for Deep Learning Based Satellite Multi-View Stereo MatchingJian Gao, Jin Liu, Shunping JiICCV 2021 · 35 citations
- Rethinking Disparity: A Depth Range Free Multi-View Stereo Based on DisparityQingsong Yan, Qiang Wang, Kaiyong Zhao, Bo Li et al.AAAI 2023 · 21 citations
- A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose EmbeddingYitong Dong, Yijin Li, Zhaoyang Huang, Weikang Bian et al.NeurIPS 2024 · 7 citations
- SemStereo: Semantic-Constrained Stereo Matching Network for Remote SensingChen Chen, Liangjin Zhao, Yuanchun He, Yingxuan Long et al.AAAI 2025 · 3 citations
- Scale Efficient Training for Large DatasetsQing Zhou, Junyu Gao, Qi WangCVPR 2025
Related papers
- BlendedMVS: A Large-Scale Dataset for Generalized Multi-View Stereo NetworksYao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang et al.CVPR 2020
- EPP-MVSNet: Epipolar-assembling based Depth Prediction for Multi-view StereoXinjun Ma, Yue Gong, Qirui Wang, Jingwei Huang et al.ICCV 2021 · 147 citations
- Multi-View Stereo Representation Revist: Region-Aware MVSNetYisu Zhang, Jianke Zhu, Lixiang LinCVPR 2023
- MVSCRF: Learning Multi-View Stereo With Conditional Random FieldsYouze Xue, Jiansheng Chen, Weitao Wan, Yiqing Huang et al.ICCV 2019 · 95 citations
- Learning Inverse Depth Regression for Multi-View Stereo with Correlation Cost VolumeQingshan Xu, Wenbing TaoAAAI 2020 · 145 citations
