RIAV-MVS: Recurrent-Indexing an Asymmetric Volume for Multi-View Stereo
Changjiang Cai, Pan Ji, Qingan Yan, Yi Xu
Abstract
This paper presents a learning-based method for multiview depth estimation from posed images. Our core idea is a "learning-to-optimize" paradigm that iteratively indexes a plane-sweeping cost volume and regresses the depth map via a convolutional Gated Recurrent Unit (GRU). Since the cost volume plays a paramount role in encoding the multiview geometry, we aim to improve its construction both at pixel-and frame-levels. At the pixel level, we propose to break the symmetry of the Siamese network (which is typically used in MVS to extract image features) by introducing a transformer block to the reference image (but not to the source images). Such an asymmetric volume allows the network to extract global features from the reference image to predict its depth map. Given potential inaccuracies in the poses between reference and source images, we propose to incorporate a residual pose network to correct the relative poses. This essentially rectifies the cost volume at the frame level. We conduct extensive experiments on real-world MVS datasets and show that our method achieves state-of-the-art performance in terms of both within-dataset evaluation and cross-dataset generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View StereoChenjie Cao, Xinlin Ren, Yanwei FuICLR 2024 · 68 citations
- One at a Time: Progressive Multi-Step Volumetric Probability Learning for Reliable 3D Scene PerceptionBohan Li, Yasheng Sun, Jingxin Dong, Zheng Zhu et al.AAAI 2024 · 9 citations
- RRT-MVS: Recurrent Regularization Transformer for Multi-View StereoJianfei Jiang, Liyong Wang, Haochen Yu, Tianyu Hu et al.AAAI 2025 · 7 citations
- UniScene: Unified Occupancy-centric Driving Scene GenerationBohan Li, Jiazhe Guo, Hongsi Liu, Yingshuang Zou et al.CVPR 2025
- Geometry Field Splatting with Gaussian SurfelsKaiwen Jiang, Venkataram Sivaram, Cheng Peng, Ravi RamamoorthiCVPR 2025
Builds on15
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 653 citations
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 403 citations
- Practical Stereo Matching via Cascaded Recurrent Network with Adaptive CorrelationJiankun Li, Peisen Wang, Pengfei Xiong, Tao Cai et al.CVPR 2022 · 294 citations
- P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Haipeng Huang et al.ICCV 2019 · 254 citations
- MonoIndoor: Towards Good Practice of Self-Supervised Monocular Depth Estimation for Indoor EnvironmentsPan Ji, Runze Li, Bir Bhanu, Yi XuICCV 2021 · 82 citations
Related papers
- Efficient Multi-view Stereo by Iterative Dynamic Cost VolumeShaoqian Wang, Bo Li, Yuchao DaiCVPR 2022 · 62 citations
- RayMVSNet: Learning Ray-based 1D Implicit Fields for Accurate Multi-View StereoJunhua Xi, Yifei Shi, Yijie Wang, Yulan Guo et al.CVPR 2022 · 129 citations
- MVSCRF: Learning Multi-View Stereo With Conditional Random FieldsYouze Xue, Jiansheng Chen, Weitao Wan, Yiqing Huang et al.ICCV 2019 · 95 citations
- RAGO: Recurrent Graph Optimizer For Multiple Rotation AveragingHeng Li, Zhaopeng Cui, Shuaicheng Liu, Ping TanCVPR 2022 · 15 citations
- DS-MVSNet: Unsupervised Multi-view Stereo via Depth SynthesisJingliang Li, Zhengda Lu, Yiqun Wang, Ying Wang et al.ACM MM 2022 · 19 citations
