TransMVSNet: Global Context-aware Multi-view Stereo Network with Transformers
Yikang Ding, Wentao Yuan, Qingtian Zhu, Haotian Zhang, Xiangyue Liu, Yuanjiang Wang, Xiao Liu
Abstract
In this paper, we present TransMVSNet, based on our exploration of feature matching in multi-view stereo (MVS). We analogize MVS back to its nature of a feature matching task and therefore propose a powerful Feature Matching Transformer (FMT) to leverage intra-(self-) and inter-(cross-) attention to aggregate long-range context information within and across images. To facilitate a better adaptation of the FMT, we leverage an Adaptive Receptive Field (ARF) module to ensure a smooth transit in scopes of features and bridge different stages with a feature pathway to pass transformed features and gradients across different scales. In addition, we apply pair-wise feature correlation to measure similarity between features, and adopt ambiguity-reducing focal loss to strengthen the supervision. To the best of our knowledge, TransMVSNet is the first attempt to leverage Transformer into the task of MVS. As a result, our method achieves state-of-the-art performance on DTU dataset, Tanks and Temples benchmark, and BlendedMVS dataset. Code is available at https: //github.com/MegviiRobot/TransMVSNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f179cddc-102d-493c-95d7-bded126d1cfeCited by top-tier papers53
- MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View StereoChenjie Cao, Xinlin Ren, Yanwei FuICLR 2024 · 68 citations
- TranSplat: Generalizable 3D Gaussian Splatting from Sparse Multi-View Images with TransformersChuanrui Zhang, Yingshuang Zou, Zhuoling Li, Minmin Yi et al.AAAI 2025 · 64 citations
- WT-MVSNet: Window-based Transformers for Multi-view StereoJinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang et al.NeurIPS 2022 · 50 citations
- When Epipolar Constraint Meets Non-local Operators in Multi-View StereoTianqi Liu, Xinyi Ye, Weiyue Zhao, Zhiyu Pan et al.ICCV 2023 · 43 citations
- Efficient Edge-Preserving Multi-View Stereo Network for Depth EstimationWanjuan Su, Wenbing TaoAAAI 2023 · 40 citations
Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding et al.ICCV 2021 · 380 citations
- AA-RMVSNet: Adaptive Aggregation Recurrent Multi-view Stereo NetworkZizhuang Wei, Qingtian Zhu, Chen Min, Yisong Chen et al.ICCV 2021 · 193 citations
Related papers
- RRT-MVS: Recurrent Regularization Transformer for Multi-View StereoJianfei Jiang, Liyong Wang, Haochen Yu, Tianyu Hu et al.AAAI 2025 · 7 citations
- Attention-Aware Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Yuesong Wang et al.CVPR 2020
- SIR-Former: Stereo Image Restoration Using TransformerZizheng Yang, Mingde Yao, Jie Huang, Man Zhou et al.ACM MM 2022 · 24 citations
- MonoMVSNet: Monocular Priors Guided Multi-View Stereo NetworkJianfei Jiang, Qiankun Liu, Haochen Yu, Hongyuan Liu et al.ICCV 2025 · 3 citations
- Stereo Video Super-Resolution via Exploiting View-Temporal CorrelationsRuikang Xu, Zeyu Xiao, Mingde Yao, Yueyi Zhang et al.ACM MM 2021 · 20 citations
