FoundationStereo: Zero-Shot Stereo Matching
Bowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz, Orazio Gallo, Stan Birchfield
Abstract
Tremendous progress has been made in deep stereo matching to excel on benchmark datasets through per-domain fine-tuning. However, achieving strong zero-shot generalization — a hallmark of foundation models in other computer vision tasks — remains challenging for stereo matching. We introduce FoundationStereo, a foundation model for stereo depth estimation designed to achieve strong zero-shot generalization. To this end, we first construct a large-scale (1M stereo pairs) synthetic training dataset featuring large diversity and high photorealism, followed by an automatic self-curation pipeline to remove ambiguous samples. We then design a number of network architecture components to enhance scalability, including a side-tuning feature backbone that adapts rich monocular priors from vision foundation models to mitigate the sim-to-real gap, and long-range context reasoning for effective cost volume filtering. Together, these components lead to strong robustness and accuracy across domains, establishing a new standard in zero-shot stereo depth estimation. Project page: https://nvlabs.github.io/FoundationStereo/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c863ffb-aa99-4957-b88d-f403b1d00d7aCited by top-tier papers52
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen et al.ICLR 2026 · 720 citations
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han et al.NeurIPS 2025 · 159 citations
- PointWorld: Scaling 3D World Models for In-The-Wild Robotic ManipulationWenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu et al.CVPR 2026 · 87 citations
- TAPIP3D: Tracking Any Point in Persistent 3D GeometryBowei Zhang, Lei Ke, Adam W. Harley, Katerina FragkiadakiNeurIPS 2025 · 79 citations
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World ModelingYang Zhou, Yifan Wang, Jianjun Zhou, Wenzheng Chang et al.ICLR 2026 · 58 citations
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
Related papers
- What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?David Yan, Alexander Raistrick, Jia DengCVPR 2026 · 8 citations
- Generalized Geometry Encoding Volume for Real-time Stereo MatchingJiaxin Liu, Gangwei Xu, Xianqi Wang, Chengliang Zhang et al.AAAI 2026
- DEFOM-Stereo: Depth Foundation Model Based Stereo MatchingHualie Jiang, Zhiqiang Lou, Laiyan Ding, Rui Xu et al.CVPR 2025
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingBowen Wen, Shaurya Dewan, Stan BirchfieldCVPR 2026 · 36 citations
- Lite Any Stereo: Efficient Zero-Shot Stereo MatchingJunpeng Jing, Weixun Luo, Ye Mao, Krystian MikolajczykCVPR 2026 · 4 citations
