A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding
Yitong Dong, Yijin Li, Zhaoyang Huang, Weikang Bian, Jingbo Liu, Hujun Bao, Zhaopeng Cui, Hongsheng Li, Guofeng Zhang
Abstract
In this paper, we propose a novel multi-view stereo (MVS) framework that gets rid of the depth range prior. Unlike recent prior-free MVS methods that work in a pair-wise manner, our method simultaneously considers all the source images. Specifically, we introduce a Multi-view Disparity Attention (MDA) module to aggregate long-range context information within and across multi-view images. Considering the asymmetry of the epipolar disparity flow, the key to our method lies in accurately modeling multi-view geometric constraints. We integrate pose embedding to encapsulate information such as multi-view camera poses, providing implicit geometric constraints for multi-view disparity feature fusion dominated by attention. Additionally, we construct corresponding hidden states for each source image due to significant differences in the observation quality of the same pixel in the reference frame across multiple source frames. We explicitly estimate the quality of the current pixel corresponding to sampled points on the epipolar line of the source image and dynamically update hidden states through the uncertainty estimation module. Extensive results on the DTU dataset and Tanks&Temple benchmark demonstrate the effectiveness of our method. The code is available at our project page: https://zju3dv.github.io/GD-PoseMVS/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4787b2b-80d2-4766-a634-6fe5c6a7f6b3Cited by top-tier papers4
- BlinkTrack: Feature Tracking Over 80 FPS via Events and ImagesYichen Shen, Yijin Li, Shuo Chen, Guanglin Li et al.ICCV 2025 · 3 citations
- One-Shot Refiner: Boosting Feed-forward Novel View Synthesis via One-Step DiffusionYitong Dong, Qi Zhang, Minchao Jiang, Zhiqiang Wu et al.AAAI 2026 · 2 citations
- Votesplat: Hough Voting Gaussian Splatting for 3D Scene UnderstandingMinchao Jiang, Shunyu Jia, Jiaming Gu, Xiaoyuan Lu et al.ICCV 2025 · 1 citation
- GeoCAD: Local Geometry-Controllable CAD Generation with Large Language ModelsZhanwei Zhang, Kaiyuan Liu, Junjie Liu, Wenxiao Wang et al.NeurIPS 2025 · 1 citation
Builds on30
- P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Haipeng Huang et al.ICCV 2019 · 254 citations
- TransMVSNet: Global Context-aware Multi-view Stereo Network with TransformersYikang Ding, Wentao Yuan, Qingtian Zhu, Haotian Zhang et al.CVPR 2022 · 236 citations
- AA-RMVSNet: Adaptive Aggregation Recurrent Multi-view Stereo NetworkZizhuang Wei, Qingtian Zhu, Chen Min, Yisong Chen et al.ICCV 2021 · 193 citations
- Rethinking Depth Estimation for Multi-View Stereo: A Unified RepresentationRui Peng, Rongjie Wang, Zhenyu Wang, Yawen Lai et al.CVPR 2022 · 159 citations
- Planar Prior Assisted PatchMatch Multi-View StereoQingshan Xu, Wenbing TaoAAAI 2020 · 154 citations
Related papers
- Rethinking Disparity: A Depth Range Free Multi-View Stereo Based on DisparityQingsong Yan, Qiang Wang, Kaiyong Zhao, Bo Li et al.AAAI 2023 · 21 citations
- Attention-Aware Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Yuesong Wang et al.CVPR 2020
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 403 citations
- Digging into Uncertainty in Self-supervised Multi-view StereoHongbin Xu, Zhipeng Zhou, Yali Wang, Wenxiong Kang et al.ICCV 2021 · 68 citations
- V-FUSE: Volumetric Depth Map Fusion with Long-Range ConstraintsNathaniel Burgdorfer, Philippos MordohaiICCV 2023 · 1 citation
