DEFOM-Stereo: Depth Foundation Model Based Stereo Matching
Hualie Jiang, Zhiqiang Lou, Laiyan Ding, Rui Xu, Minglang Tan, Wenjie Jiang, Rui Huang
摘要
Stereo matching is a key technique for metric depth estimation in computer vision and robotics. Real-world challenges like occlusion and non-texture hinder accurate disparity estimation from binocular matching cues. Recently, monocular relative depth estimation has shown remarkable generalization using vision foundation models. Thus, to facilitate robust stereo matching with monocular depth cues, we incorporate a robust monocular relative depth model into the recurrent stereo-matching framework, building a new framework for depth foundation model-based stereomatching, DEFOM-Stereo. In the feature extraction stage, we construct the combined context and matching feature encoder by integrating features from conventional CNNs and DEFOM. In the update stage, we use the depth predicted by DEFOM to initialize the recurrent disparity and introduce a scale update module to refine the disparity at the correct scale. DEFOM-Stereo is verified to have much stronger zero-shot generalization compared with SOTA methods. Moreover, DEFOM-Stereo achieves top performance on the KITTI 2012, KITTI 2015, Middlebury, and ETH3D benchmarks, ranking 1 st on many metrics. In the joint evaluation under the robust vision challenge, our model simultaneously outperforms previous models on the individual benchmarks, further demonstrating its outstanding capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingBowen Wen, Shaurya Dewan, Stan BirchfieldCVPR 2026 · 被引用 36 次
- Large Depth Completion Model from Sparse ObservationsZhu Yu, zhengyi zhao, Runmin Zhang, Lingteng Qiu 等ICLR 2026 · 被引用 8 次
- S2M2: Scalable Stereo Matching Model for Reliable Depth EstimationJunhong Min, Youngpil Jeon, Jimin Kim, Minyong ChoiICCV 2025 · 被引用 8 次
- BANet: Bilateral Aggregation Network for Mobile Stereo MatchingGangwei Xu, Jiaxin Liu, Xianqi Wang, Junda Cheng 等ICCV 2025 · 被引用 7 次
- BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent AlignmentTongfan Guan, Jiaxin Guo, Chen Wang, Yun-Hui LiuICCV 2025 · 被引用 6 次
它引用的顶会 Paper29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding 等ICCV 2021 · 被引用 380 次
相关 Paper
- Generalized Geometry Encoding Volume for Real-time Stereo MatchingJiaxin Liu, Gangwei Xu, Xianqi Wang, Chengliang Zhang 等AAAI 2026
- FoundationStereo: Zero-Shot Stereo MatchingBowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz 等CVPR 2025
- MonSter: Marry Monodepth to Stereo Unleashes PowerJunda Cheng, Longliang Liu, Gangwei Xu, Xianqi Wang 等CVPR 2025
- PromptStereo: Zero-Shot Stereo Matching via Structure and Motion PromptsXianqi Wang, Hao Yang, Hangtian Wang, JunDa Cheng 等CVPR 2026 · 被引用 5 次
- MonoMVSNet: Monocular Priors Guided Multi-View Stereo NetworkJianfei Jiang, Qiankun Liu, Haochen Yu, Hongyuan Liu 等ICCV 2025 · 被引用 3 次
