A Recipe for Generating 3D Worlds from a Single Image
Katja Schwarz, Denis Rozumny, Samuel Rota Bulò, Lorenzo Porzi, Peter Kontschieder
Abstract
We introduce a recipe for generating immersive 3D worlds from a single image by framing the task as an in-context learning problem for 2D inpainting models. This approach requires minimal training and uses existing generative models. Our process involves two steps: generating coherent panoramas using a pre-trained diffusion model and lifting these into 3D with a metric depth estimator. We then fill unobserved regions by conditioning the inpainting model on rendered point clouds, requiring minimal fine-tuning. Tested on both synthetic and real images, our method produces high-quality 3D environments suitable for VR display. By explicitly modeling the 3D structure of the generated environment from the start, our approach consistently outperforms state-of-the-art, video synthesis-based methods along multiple quantitative image quality metrics. Project Page: https://katjaschwarz.github.io/worlds/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6a5a594-eab3-4fae-9712-19a079d04f82Cited by top-tier papers6
- Sharp Monocular View Synthesis in Less Than a SecondLars Mescheder, Wei Dong, Shiwei Li, Xuyang Bai et al.ICLR 2026 · 25 citations
- Gen3R: 3D Scene Generation Meets Feed-Forward ReconstructionJiaxin Huang, Yuanbo Yang, Bangbang Yang, Lin Ma et al.CVPR 2026 · 24 citations
- SphericalDreamer: Generating Navigable Immersive 3D Worlds with Panorama FusionAntoine Schnepf, Karim Kassab, Flavian Vasile, Andrew ComportICML 2026 · 1 citation
- MiDSummer: Multi-Guidance Diffusion for Controllable Zero-Shot Immersive Gaussian Splatting Scene GenerationAnjun Hu, Richard Tomsett, Valentin Gourmet, Massimo Camplani et al.ICCV 2025
- PrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic Urban ScenesChristina Ourania Tze, Daniel Dauner, Yiyi Liao, Dzmitry Tsishkou et al.CVPR 2026
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single ImagesPhilipp Wulff, Felix Wimbauer, Dominik Muhle, Daniel CremersICCV 2025 · 1 citation
- Vistadream: Sampling Multiview Consistent Images for Single-View Scene ReconstructionHaiping Wang, Yuan Liu, Ziwei Liu, Wenping Wang et al.ICCV 2025 · 8 citations
- 3D-aware Image Generation using 2D Diffusion ModelsJianfeng Xiang, Jiaolong Yang, Binbin Huang, Xin TongICCV 2023 · 82 citations
- SynSin: End-to-End View Synthesis From a Single ImageOlivia Wiles, Georgia Gkioxari, Richard Szeliski, Justin JohnsonCVPR 2020
- Make-It-4D: Synthesizing a Consistent Long-Term Dynamic Scene Video from a Single ImageLiao Shen, Xingyi Li, Huiqiang Sun, Juewen Peng et al.ACM MM 2023 · 15 citations
