ZeroStereo: Zero-Shot Stereo Matching from Single Images
Xianqi Wang, Hao Yang, Gangwei Xu, Junda Cheng, Min Lin, Yong Deng, Jinliang Zang, Yurui Chen, Xin Yang
Abstract
State-of-the-art supervised stereo matching methods have achieved remarkable performance on various benchmarks. However, their generalization to real-world scenarios remains challenging due to the scarcity of annotated real-world stereo data. In this paper, we propose ZeroStereo, a novel stereo image generation pipeline for zero-shot stereo matching. Our approach synthesizes high-quality right images from arbitrary single images by leveraging pseudo disparities generated by a monocular depth estimation model. Unlike previous methods that address occluded regions by filling missing areas with neighboring pixels or random backgrounds, we fine-tune a diffusion inpainting model to recover missing details while preserving semantic structure. Additionally, we propose Training-Free Confidence Generation, which mitigates the impact of unreliable pseudo labels without additional training, and Adaptive Disparity Selection, which ensures a diverse and realistic disparity distribution while preventing excessive occlusion and foreground distortion. Experiments demonstrate that models trained with our pipeline achieve state-of-theart zero-shot generalization across multiple datasets, with only a dataset volume comparable to Scene Flow. Code: https://github.com/Windsrain/ZeroStereo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed9124eb-bc81-4a76-8f66-4b7114310a05Cited by top-tier papers7
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingBowen Wen, Shaurya Dewan, Stan BirchfieldCVPR 2026 · 36 citations
- Jasmine: Harnessing Diffusion Prior for Self-supervised Depth EstimationJiyuan Wang, Chunyu Lin, Cheng Guan, Lang Nie et al.NeurIPS 2025 · 26 citations
- BANet: Bilateral Aggregation Network for Mobile Stereo MatchingGangwei Xu, Jiaxin Liu, Xianqi Wang, Junda Cheng et al.ICCV 2025 · 7 citations
- PromptStereo: Zero-Shot Stereo Matching via Structure and Motion PromptsXianqi Wang, Hao Yang, Hangtian Wang, JunDa Cheng et al.CVPR 2026 · 5 citations
- Pip-Stereo: Progressive Iterations Pruner for Iterative Optimization based Stereo MatchingJintu Zheng, Qizhe Liu, Huangxin Xu, zhuojie ChenCVPR 2026 · 1 citation
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
Related papers
- RobuSTereo: Robust Zero-Shot Stereo Matching under Adverse WeatherYuran Wang, Yingping Liang, Yutao Hu, Ying FuICCV 2025 · 3 citations
- FoundationStereo: Zero-Shot Stereo MatchingBowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz et al.CVPR 2025
- Towards Open-World Generation of Stereo Images and Unsupervised MatchingFeng Qiao, Zhexiao Xiong, Eric Xing, Nathan JacobsICCV 2025 · 1 citation
- NeRF-Supervised Deep StereoFabio Tosi, Alessio Tonioni, Daniele De Gregorio, Matteo PoggiCVPR 2023
- DepthFM: Fast Generative Monocular Depth Estimation with Flow MatchingMing Gui, Johannes Schusterbauer, Ulrich Prestel, Pingchuan Ma et al.AAAI 2025 · 51 citations
