Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
Bowen Wen, Shaurya Dewan, Stan Birchfield
Abstract
Stereo foundation models achieve strong zero-shotgeneralization but remain computationally prohibitive forreal-time applications. Efficient stereo architectures, on the other hand, sacrificerobustness for speed and require costly per-domain fine-tuning.To bridge this gap, we present Fast-FoundationStereo, a family of architectures that achieve, for the first time, strong zero-shot generalization at real-time frame rate. We employ a divide-and-conquer acceleration strategy with three components: (1) knowledge distillation to compress the hybrid backbone into a single efficient student; (2) blockwise neural architecture search for automatically discovering optimal cost filtering designs under latency budgets, reducing search complexity exponentially; and (3) structured pruning for eliminating redundancy in the iterative refinement module. Furthermore, we introduce an automatic pseudo-labeling pipeline used to curate 1.4M in-the-wild stereo pairs to supplement synthetic training data and facilitate knowledge distillation. The resulting model can run over 10× faster than FoundationStereo while closely matching its zero-shot accuracy, thus establishing a new state-of-the-art among real-time methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a49b5626-fce7-44cb-9799-222803c3aa44Builds on37
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Attention Concatenation Volume for Accurate and Efficient Stereo MatchingGangwei Xu, Junda Cheng, Peng Guo, Xin YangCVPR 2022 · 265 citations
- Compact Language Models via Pruning and Knowledge DistillationSaurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski et al.NeurIPS 2024 · 198 citations
Related papers
- FoundationStereo: Zero-Shot Stereo MatchingBowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz et al.CVPR 2025
- Lite Any Stereo: Efficient Zero-Shot Stereo MatchingJunpeng Jing, Weixun Luo, Ye Mao, Krystian MikolajczykCVPR 2026 · 4 citations
- Generalized Geometry Encoding Volume for Real-time Stereo MatchingJiaxin Liu, Gangwei Xu, Xianqi Wang, Chengliang Zhang et al.AAAI 2026
- DeepPruner: Learning Efficient Stereo Matching via Differentiable PatchMatchShivam Duggal, Shenlong Wang, Wei-Chiu Ma, Rui Hu et al.ICCV 2019 · 300 citations
- PromptStereo: Zero-Shot Stereo Matching via Structure and Motion PromptsXianqi Wang, Hao Yang, Hangtian Wang, JunDa Cheng et al.CVPR 2026 · 5 citations
