RIAV-MVS: Recurrent-Indexing an Asymmetric Volume for Multi-View Stereo
Changjiang Cai, Pan Ji, Qingan Yan, Yi Xu
摘要
This paper presents a learning-based method for multiview depth estimation from posed images. Our core idea is a "learning-to-optimize" paradigm that iteratively indexes a plane-sweeping cost volume and regresses the depth map via a convolutional Gated Recurrent Unit (GRU). Since the cost volume plays a paramount role in encoding the multiview geometry, we aim to improve its construction both at pixel-and frame-levels. At the pixel level, we propose to break the symmetry of the Siamese network (which is typically used in MVS to extract image features) by introducing a transformer block to the reference image (but not to the source images). Such an asymmetric volume allows the network to extract global features from the reference image to predict its depth map. Given potential inaccuracies in the poses between reference and source images, we propose to incorporate a residual pose network to correct the relative poses. This essentially rectifies the cost volume at the frame level. We conduct extensive experiments on real-world MVS datasets and show that our method achieves state-of-the-art performance in terms of both within-dataset evaluation and cross-dataset generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View StereoChenjie Cao, Xinlin Ren, Yanwei FuICLR 2024 · 被引用 68 次
- One at a Time: Progressive Multi-Step Volumetric Probability Learning for Reliable 3D Scene PerceptionBohan Li, Yasheng Sun, Jingxin Dong, Zheng Zhu 等AAAI 2024 · 被引用 9 次
- RRT-MVS: Recurrent Regularization Transformer for Multi-View StereoJianfei Jiang, Liyong Wang, Haochen Yu, Tianyu Hu 等AAAI 2025 · 被引用 7 次
- UniScene: Unified Occupancy-centric Driving Scene GenerationBohan Li, Jiazhe Guo, Hongsi Liu, Yingshuang Zou 等CVPR 2025
- Geometry Field Splatting with Gaussian SurfelsKaiwen Jiang, Venkataram Sivaram, Cheng Peng, Ravi RamamoorthiCVPR 2025
它引用的顶会 Paper15
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 被引用 653 次
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 被引用 403 次
- Practical Stereo Matching via Cascaded Recurrent Network with Adaptive CorrelationJiankun Li, Peisen Wang, Pengfei Xiong, Tao Cai 等CVPR 2022 · 被引用 294 次
- P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Haipeng Huang 等ICCV 2019 · 被引用 254 次
- MonoIndoor: Towards Good Practice of Self-Supervised Monocular Depth Estimation for Indoor EnvironmentsPan Ji, Runze Li, Bir Bhanu, Yi XuICCV 2021 · 被引用 82 次
相关 Paper
- Efficient Multi-view Stereo by Iterative Dynamic Cost VolumeShaoqian Wang, Bo Li, Yuchao DaiCVPR 2022 · 被引用 62 次
- RayMVSNet: Learning Ray-based 1D Implicit Fields for Accurate Multi-View StereoJunhua Xi, Yifei Shi, Yijie Wang, Yulan Guo 等CVPR 2022 · 被引用 129 次
- MVSCRF: Learning Multi-View Stereo With Conditional Random FieldsYouze Xue, Jiansheng Chen, Weitao Wan, Yiqing Huang 等ICCV 2019 · 被引用 95 次
- RAGO: Recurrent Graph Optimizer For Multiple Rotation AveragingHeng Li, Zhaopeng Cui, Shuaicheng Liu, Ping TanCVPR 2022 · 被引用 15 次
- DS-MVSNet: Unsupervised Multi-view Stereo via Depth SynthesisJingliang Li, Zhengda Lu, Yiqun Wang, Ying Wang 等ACM MM 2022 · 被引用 19 次
