Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching
Xiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai, Feitong Tan, Ping Tan
Abstract
The deep multi-view stereo (MVS) and stereo matching approaches generally construct 3D cost volumes to regularize and regress the output depth or disparity. These methods are limited when high-resolution outputs are needed since the memory and time costs grow cubically as the volume resolution increases. In this paper, we propose a both memory and time efficient cost volume formulation that is complementary to existing multi-view stereo and stereo matching approaches based on 3D cost volumes. First, the proposed cost volume is built upon a standard feature pyramid encoding geometry and context at gradually finer scales. Then, we can narrow the depth (or disparity) range of each stage by the depth (or disparity) map from the previous stage. With gradually higher cost volume resolution and adaptive adjustment of depth (or disparity) intervals, the output is recovered in a coarser to fine manner. We apply the cascade cost volume to the representative MVS-Net, and obtain a 35.6% improvement on DTU benchmark (1st place), with 50.6% and 59.3% reduction in GPU memory and run-time. It is also the state-of-the-art learning-based method on Tanks and Temples benchmark. The statistics of accuracy, run-time and GPU memory on other representative stereo CNNs also validate the effectiveness of our proposed method. Our source code is available at https://github.com/alibaba/cascade-stereo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers204
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang et al.ICCV 2021 · 1,024 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding et al.ICCV 2021 · 380 citations
- Neural 3D Video Synthesis from Multi-view VideoTianye Li, Mira Slavcheva, Michael Zollhöfer, Simon Green et al.CVPR 2022 · 324 citations
- Attention Concatenation Volume for Accurate and Efficient Stereo MatchingGangwei Xu, Junda Cheng, Peng Guo, Xin YangCVPR 2022 · 265 citations
Builds on5
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 403 citations
- P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Haipeng Huang et al.ICCV 2019 · 254 citations
- Semantic Stereo Matching With Pyramid Cost VolumesZhenyao Wu, Xinyi Wu, Xiaoping Zhang, Song Wang et al.ICCV 2019 · 125 citations
- Multi-View Stereo by Temporal Nonparametric FusionYuxin Hou, Juho Kannala, Arno SolinICCV 2019 · 99 citations
- TAPA-MVS: Textureless-Aware PAtchMatch Multi-View StereoAndrea Romanoni, Matteo MatteucciICCV 2019 · 95 citations
Related papers
- Efficient Multi-view Stereo by Iterative Dynamic Cost VolumeShaoqian Wang, Bo Li, Yuchao DaiCVPR 2022 · 62 citations
- Learning Inverse Depth Regression for Multi-View Stereo with Correlation Cost VolumeQingshan Xu, Wenbing TaoAAAI 2020 · 145 citations
- Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton RefinementZehao Yu, Shenghua GaoCVPR 2020
- PatchmatchNet: Learned Multi-View Patchmatch StereoFangjinhua Wang, Silvano Galliani, Christoph Vogel, Pablo Speciale et al.CVPR 2021
- Generalized Binary Search Network for Highly-Efficient Multi-View StereoZhenxing Mi, Di Chang, Dan XuCVPR 2022
