MVSAnywhere: Zero-Shot Multi-View Stereo
Sergio Izquierdo, Mohamed Sayed, Michael Firman, Guillermo Garcia-Hernando, Daniyar Turmukhambetov, Javier Civera, Oisin Mac Aodha, Gabriel J. Brostow, Jamie Watson
摘要
Computing accurate depth from multiple views is a fundamental and longstanding challenge in computer vision. However, most existing approaches do not generalize well across different domains and scene types (e.g. indoor vs. outdoor). Training a general-purpose multi-view stereo model is challenging and raises several questions, e.g. how to best make use of transformer-based architectures, how to incorporate additional metadata when there is a variable number of input views, and how to estimate the range of valid depths which can vary considerably across different scenes and is typically not known a priori? To address these issues, we introduce MVSA, a novel and versatile Multi-View Stereo architecture that aims to work Anywhere by generalizing across diverse domains and depth ranges. MVSA combines monocular and multi-view cues with an adaptive cost volume to deal with scale-related issues. We demonstrate state-of-the-art zero-shot depth estimation on the Robust Multi-View Depth Benchmark, surpassing existing multi-view stereo and monocular baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface ReconstructionJiahe Li, Jiawei Zhang, Youmin Zhang, Xiao Bai 等NeurIPS 2025 · 被引用 18 次
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with BackendHengyi Wang, Lourdes AgapitoCVPR 2026 · 被引用 17 次
- SparseSplat: Towards Applicable Feed-Forward 3D Gaussian Splatting with Pixel-Unaligned PredictionZicheng Zhang, Xiangting Meng, Ke Wu, Wenchao DingCVPR 2026 · 被引用 7 次
- AnchorSplat: Feed-Forward 3D Gaussian Splatting With 3D Geometric PriorsXiaoxue Zhang, Xiaoxu Zheng, Yixuan Yin, Tiao Zhao 等CVPR 2026 · 被引用 6 次
- FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion BasesMatteo Poggi, Fabio TosiICCV 2025 · 被引用 5 次
它引用的顶会 Paper42
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar 等ICCV 2021 · 被引用 633 次
相关 Paper
- FoundationStereo: Zero-Shot Stereo MatchingBowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz 等CVPR 2025
- RRT-MVS: Recurrent Regularization Transformer for Multi-View StereoJianfei Jiang, Liyong Wang, Haochen Yu, Tianyu Hu 等AAAI 2025 · 被引用 7 次
- MonoMVSNet: Monocular Priors Guided Multi-View Stereo NetworkJianfei Jiang, Qiankun Liu, Haochen Yu, Hongyuan Liu 等ICCV 2025 · 被引用 3 次
- UniDepth: Universal Monocular Metric Depth EstimationLuigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Segù 等CVPR 2024 · 被引用 122 次
- LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous DrivingQihao Sun, Jiarun Liu, Ziqian Ni, Jianyun Xu 等CVPR 2026 · 被引用 2 次
