CL-MVSNet: Unsupervised Multi-view Stereo with Dual-level Contrastive Learning
Kaiqiang Xiong, Rui Peng, Zhe Zhang, Tianxing Feng, Jianbo Jiao, Feng Gao, Ronggang Wang
Abstract
Unsupervised Multi-View Stereo (MVS) methods have achieved promising progress recently. However, previous methods primarily depend on the photometric consistency assumption, which may suffer from two limitations: indistinguishable regions and view-dependent effects, e.g., low-textured areas and reflections. To address these issues, in this paper, we propose a new dual-level contrastive learning approach, named CL-MVSNet. Specifically, our model integrates two contrastive branches into an unsupervised MVS framework to construct additional supervisory signals. On the one hand, we present an image-level contrastive branch to guide the model to acquire more context awareness, thus leading to more complete depth estimation in indistinguishable regions. On the other hand, we exploit a scene-level contrastive branch to boost the representation ability, improving robustness to view-dependent effects. Moreover, to recover more accurate 3D geometry, we introduce an ℒ0.5 photometric consistency loss, which encourages the model to focus more on accurate points while mitigating the gradient penalty of undesirable ones. Extensive experiments on DTU and Tanks&Temples benchmarks demonstrate that our approach achieves state-of-the-art performance among all end-to-end unsupervised MVS frameworks and outperforms its supervised counterpart by a considerable margin without fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20eb8a6d-e499-4282-94f6-4256200326d1Cited by top-tier papers5
- 4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic ScenesJinbo Yan, Rui Peng, Luyang Tang, Ronggang WangACM MM 2024 · 23 citations
- ClipGStream: Clip-Stream Gaussian Splatting for Any Length and Any Motion Multi-View Dynamic Scene ReconstructionJie Liang, Jiahao Wu, Chao Wang, Jiayu Yang et al.CVPR 2026 · 3 citations
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D ReconstructionHeng Jia, Linchao Zhu, Na ZhaoICCV 2025 · 1 citation
- Sparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse ViewsJiang Wu, Rui Li, Yu Zhu, Rong Guo et al.CVPR 2025
- Geometry Field Splatting with Gaussian SurfelsKaiwen Jiang, Venkataram Sivaram, Cheng Peng, Ravi RamamoorthiCVPR 2025
Builds on29
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View StereoKeyang Luo, Tao Guan, Lili Ju, Haipeng Huang et al.ICCV 2019 · 254 citations
- TransMVSNet: Global Context-aware Multi-view Stereo Network with TransformersYikang Ding, Wentao Yuan, Qingtian Zhu, Haotian Zhang et al.CVPR 2022 · 236 citations
- Rethinking Depth Estimation for Multi-View Stereo: A Unified RepresentationRui Peng, Rongjie Wang, Zhenyu Wang, Yawen Lai et al.CVPR 2022 · 159 citations
Related papers
- Self-supervised Multi-view Stereo via Inter and Intra Network Pseudo DepthKe Qiu, Yawen Lai, Shiyi Liu, Ronggang WangACM MM 2022 · 9 citations
- MonoMVSNet: Monocular Priors Guided Multi-View Stereo NetworkJianfei Jiang, Qiankun Liu, Haochen Yu, Hongyuan Liu et al.ICCV 2025 · 3 citations
- DualNet: Robust Self-Supervised Stereo Matching with Pseudo-Label SupervisionYun Wang, Jiahao Zheng, Chenghao Zhang, Zhanjie Zhang et al.AAAI 2025 · 12 citations
- Self-Supervised Multi-view Stereo via Adjacent Geometry Guided Volume CompletionLuoyuan Xu, Tao Guan, Yuesong Wang, Yawei Luo et al.ACM MM 2022 · 16 citations
- Multi-View Stereo Representation Revist: Region-Aware MVSNetYisu Zhang, Jianke Zhu, Lixiang LinCVPR 2023
