Rethinking Depth Estimation for Multi-View Stereo: A Unified Representation
Rui Peng, Rongjie Wang, Zhenyu Wang, Yawen Lai, Ronggang Wang
摘要
Depth estimation is solved as a regression or classification problem in existing learning-based multi-view stereo methods. Although these two representations have recently demonstrated their excellent performance, they still have apparent shortcomings, e.g., regression methods tend to overfit due to the indirect learning cost volume, and classification methods cannot directly infer the exact depth due to its discrete prediction. In this paper, we propose a novel representation, termed Unification, to unify the advantages of regression and classification. It can directly constrain the cost volume like classification methods, but also realize the sub-pixel depth prediction like regression methods. To excavate the potential of unification, we design a new loss function named Unified Focal Loss, which is more uniform and reasonable to combat the challenge of sample imbalance. Combining these two unburdened modules, we present a coarse-to-fine framework, that we call UniMVSNet. The results of ranking first on both DTU and Tanks and Temples benchmarks verify that our model not only performs the best but also has the best generalization ability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View StereoChenjie Cao, Xinlin Ren, Yanwei FuICLR 2024 · 被引用 68 次
- TranSplat: Generalizable 3D Gaussian Splatting from Sparse Multi-View Images with TransformersChuanrui Zhang, Yingshuang Zou, Zhuoling Li, Minmin Yi 等AAAI 2025 · 被引用 64 次
- WT-MVSNet: Window-based Transformers for Multi-view StereoJinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang 等NeurIPS 2022 · 被引用 50 次
- When Epipolar Constraint Meets Non-local Operators in Multi-View StereoTianqi Liu, Xinyi Ye, Weiyue Zhao, Zhiyu Pan 等ICCV 2023 · 被引用 43 次
它引用的顶会 Paper14
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen 等NeurIPS 2020 · 被引用 2,118 次
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 被引用 403 次
- P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Haipeng Huang 等ICCV 2019 · 被引用 254 次
- Adaptive Unimodal Cost Volume Filtering for Deep Stereo MatchingYoumin Zhang, Yimin Chen, Xiao Bai, Suihanjin Yu 等AAAI 2020 · 被引用 201 次
- AA-RMVSNet: Adaptive Aggregation Recurrent Multi-view Stereo NetworkZizhuang Wei, Qingtian Zhu, Chen Min, Yisong Chen 等ICCV 2021 · 被引用 193 次
相关 Paper
- Learning Inverse Depth Regression for Multi-View Stereo with Correlation Cost VolumeQingshan Xu, Wenbing TaoAAAI 2020 · 被引用 145 次
- CL-MVSNet: Unsupervised Multi-view Stereo with Dual-level Contrastive LearningKaiqiang Xiong, Rui Peng, Zhe Zhang, Tianxing Feng 等ICCV 2023 · 被引用 24 次
- Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton RefinementZehao Yu, Shenghua GaoCVPR 2020
- MonoMVSNet: Monocular Priors Guided Multi-View Stereo NetworkJianfei Jiang, Qiankun Liu, Haochen Yu, Hongyuan Liu 等ICCV 2025 · 被引用 3 次
- Constraining Depth Map Geometry for Multi-View Stereo: A Dual-Depth Approach with Saddle-shaped Depth CellsXinyi Ye, Weiyue Zhao, Tianqi Liu, Zihao Huang 等ICCV 2023 · 被引用 29 次
