DEFOM-Stereo: Depth Foundation Model Based Stereo Matching
Hualie Jiang, Zhiqiang Lou, Laiyan Ding, Rui Xu, Minglang Tan, Wenjie Jiang, Rui Huang
Abstract
Stereo matching is a key technique for metric depth estimation in computer vision and robotics. Real-world challenges like occlusion and non-texture hinder accurate disparity estimation from binocular matching cues. Recently, monocular relative depth estimation has shown remarkable generalization using vision foundation models. Thus, to facilitate robust stereo matching with monocular depth cues, we incorporate a robust monocular relative depth model into the recurrent stereo-matching framework, building a new framework for depth foundation model-based stereomatching, DEFOM-Stereo. In the feature extraction stage, we construct the combined context and matching feature encoder by integrating features from conventional CNNs and DEFOM. In the update stage, we use the depth predicted by DEFOM to initialize the recurrent disparity and introduce a scale update module to refine the disparity at the correct scale. DEFOM-Stereo is verified to have much stronger zero-shot generalization compared with SOTA methods. Moreover, DEFOM-Stereo achieves top performance on the KITTI 2012, KITTI 2015, Middlebury, and ETH3D benchmarks, ranking 1 st on many metrics. In the joint evaluation under the robust vision challenge, our model simultaneously outperforms previous models on the individual benchmarks, further demonstrating its outstanding capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4311e7e-165d-46de-af14-988c40f731efCited by top-tier papers20
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingBowen Wen, Shaurya Dewan, Stan BirchfieldCVPR 2026 · 36 citations
- Large Depth Completion Model from Sparse ObservationsZhu Yu, zhengyi zhao, Runmin Zhang, Lingteng Qiu et al.ICLR 2026 · 8 citations
- S2M2: Scalable Stereo Matching Model for Reliable Depth EstimationJunhong Min, Youngpil Jeon, Jimin Kim, Minyong ChoiICCV 2025 · 8 citations
- BANet: Bilateral Aggregation Network for Mobile Stereo MatchingGangwei Xu, Jiaxin Liu, Xianqi Wang, Junda Cheng et al.ICCV 2025 · 7 citations
- BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent AlignmentTongfan Guan, Jiaxin Guo, Chen Wang, Yun-Hui LiuICCV 2025 · 6 citations
Builds on29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding et al.ICCV 2021 · 380 citations
Related papers
- Generalized Geometry Encoding Volume for Real-time Stereo MatchingJiaxin Liu, Gangwei Xu, Xianqi Wang, Chengliang Zhang et al.AAAI 2026
- FoundationStereo: Zero-Shot Stereo MatchingBowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz et al.CVPR 2025
- MonSter: Marry Monodepth to Stereo Unleashes PowerJunda Cheng, Longliang Liu, Gangwei Xu, Xianqi Wang et al.CVPR 2025
- PromptStereo: Zero-Shot Stereo Matching via Structure and Motion PromptsXianqi Wang, Hao Yang, Hangtian Wang, JunDa Cheng et al.CVPR 2026 · 5 citations
- MonoMVSNet: Monocular Priors Guided Multi-View Stereo NetworkJianfei Jiang, Qiankun Liu, Haochen Yu, Hongyuan Liu et al.ICCV 2025 · 3 citations
