Towards Open-World Generation of Stereo Images and Unsupervised Matching
Feng Qiao, Zhexiao Xiong, Eric Xing, Nathan Jacobs
Abstract
Stereo images are fundamental to numerous applications, including extended reality () devices, autonomous driving, and robotics. Unfortunately, acquiring highquality stereo images remains challenging due to the precise calibration requirements of dual-camera setups and the complexity of obtaining accurate, dense disparity maps. Existing stereo image generation methods typically focus on either visual quality for viewing or geometric accuracy for matching, but not both. We introduce GenStereo, a diffusion-based approach, to bridge this gap. The method includes two primary innovations (1) conditioning the diffusion process on a disparity-aware coordinate embedding and a warped input image, allowing for more precise stereo alignment than previous methods, and (2) an adaptive fusion mechanism that intelligently combines the diffusiongenerated image with a warped image, improving both realism and disparity consistency. Through extensive training on 11 diverse stereo datasets, GenStereo demonstrates strong generalization ability. GenStereo achieves state-of-the-art performance in both stereo image generation and unsupervised stereo matching tasks. Project page is available at https://qjizhi.github.io/genstereo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- StereoWorld: Geometry-Aware Monocular-to-Stereo Video GenerationKe Xing, Longfei Li, Yuyang Yin, Hanwen Liang et al.CVPR 2026 · 3 citations
- Elastic3D: Controllable Stereo Video Conversion with Guided Latent DecodingNando Metzger, Prune Truong, Goutam Bhat, Konrad Schindler et al.CVPR 2026 · 3 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- ZeroStereo: Zero-Shot Stereo Matching from Single ImagesXianqi Wang, Hao Yang, Gangwei Xu, Junda Cheng et al.ICCV 2025 · 1 citation
- DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video GenerationJian Shi, Qian Wang, Zhenyu Li, Wenqing Cui et al.SIGGRAPH 2026
- SDiD:Shared diffusion prior for efficient distributed stereo image compressionYichong Xia, Yimin Zhou, Zongyu Li, Shiyu Qin et al.ICML 2026
- RobuSTereo: Robust Zero-Shot Stereo Matching under Adverse WeatherYuran Wang, Yingping Liang, Yutao Hu, Ying FuICCV 2025 · 3 citations
- 3D-aware Image Generation using 2D Diffusion ModelsJianfeng Xiang, Jiaolong Yang, Binbin Huang, Xin TongICCV 2023 · 82 citations
