Lune

CVPR2026Top-tier venue

SPE-MVS: Spatial Position Encoding Enhanced Multi-View Stereo with Monocular Depth Priors

Shaoqian Wang, Jiadai Sun, Bosen Hou, Qiang Wang, Bin Fan, Bo Liu, Bin Lu, Yuchao Dai

2026Year

Abstract

Learning-based Multi-View Stereo (MVS) methods have become the mainstream in the field, relying on the construction of cost volumes through multi-view feature similarity computation. However, existing methods depend heavily on photometric consistency across views, leading to poor performance in challenging regions. To overcome this limitation, we propose SPE-MVS, a novel MVS framework enhanced with Spatial Position Encoding (SPE). The SPE represents the 3D positional information of pixels in each image within a unified metric space, constructed using monocular depth priors. We integrate the SPE alongside image data as input and introduce a Photometric-Spatial Hybrid Feature Extractor, along with an SPE-enhanced cost volume construction module. These components incorporate spatial position-based similarity computation, substantially improving robustness in challenging areas. Furthermore, we propose a Monocular Depth-guided Enhancement (MDGE) module that enhances depth probability map using monocular depth priors, thereby further boosting the depth estimation performance. Extensive experiments demonstrate that our method significantly improves reconstruction quality in difficult regions and achieves stateof-the-art (SOTA) performance on multiple benchmarks. The code will be released at https://github.com/ bdwsq1996/SPE-MVS.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ba37fc71-84a8-4f08-b8f5-75e6dbd68805

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines