WT-MVSNet: Window-based Transformers for Multi-view Stereo
Jinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang, Shihao Ren, Jia Guo, Wensen Feng, Kai Zhang
摘要
Recently, Transformers were shown to enhance the performance of multi-view stereo by enabling long-range feature interaction. In this work, we propose Window-based Transformers (WT) for local feature matching and global feature aggregation in multi-view stereo. We introduce a Window-based Epipolar Transformer (WET) which reduces matching redundancy by using epipolar constraints. Since point-to-line matching is sensitive to erroneous camera pose and calibration, we match windows near the epipolar lines. A second Shifted WT is employed for aggregating global information within cost volume. We present a novel Cost Transformer (CT) to replace 3D convolutions for cost volume regularization. In order to better constrain the estimated depth maps from multiple views, we further design a novel geometric consistency loss (Geo Loss) which punishes unreliable areas where multi-view consistency is not satisfied. Our WT multi-view stereo method (WT-MVSNet) achieves state-of-the-art performance across multiple datasets and ranks on Tanks and Temples benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View StereoChenjie Cao, Xinlin Ren, Yanwei FuICLR 2024 · 被引用 68 次
- GoMVS: Geometrically Consistent Cost Aggregation for Multi-View StereoJiang Wu, Rui Li, Haofei Xu, Wenxun Zhao 等CVPR 2024 · 被引用 34 次
- C2F2NeUS: Cascade Cost Frustum Fusion for High Fidelity and Generalizable Neural Surface ReconstructionLuoyuan Xu, Tao Guan, Yuesong Wang, Wenkai Liu 等ICCV 2023 · 被引用 25 次
- RRT-MVS: Recurrent Regularization Transformer for Multi-View StereoJianfei Jiang, Liyong Wang, Haochen Yu, Tianyu Hu 等AAAI 2025 · 被引用 7 次
- MVSMamba: Multi-View Stereo with State Space ModelJianfei Jiang, Qiankun Liu, Hongyuan Liu, Haochen Yu 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped WindowsXiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang 等CVPR 2022 · 被引用 1,207 次
相关 Paper
- When Epipolar Constraint Meets Non-local Operators in Multi-View StereoTianqi Liu, Xinyi Ye, Weiyue Zhao, Zhiyu Pan 等ICCV 2023 · 被引用 43 次
- TransMVSNet: Global Context-aware Multi-view Stereo Network with TransformersYikang Ding, Wentao Yuan, Qingtian Zhu, Haotian Zhang 等CVPR 2022 · 被引用 236 次
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding 等ICCV 2021 · 被引用 380 次
- S2M2: Scalable Stereo Matching Model for Reliable Depth EstimationJunhong Min, Youngpil Jeon, Jimin Kim, Minyong ChoiICCV 2025 · 被引用 8 次
- CATs: Cost Aggregation Transformers for Visual CorrespondenceSeokju Cho, Sunghwan Hong, Sangryul Jeon, Yunsung Lee 等NeurIPS 2021 · 被引用 133 次
