HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Models
Yiwen Chen, Hieu T. Nguyen, Vikram Voleti, Varun Jampani, Huaizu Jiang
Abstract
We introduce HouseCrafter, a novel approach that can lift a 2D floorplan into a complete large 3D indoor scene (e.g., a house). Our key insight is to adapt a 2D diffusion model, which is trained on web-scale images, to generate consistent multi-view color () and depth () images across different locations of the scene. Specifically, the RGB-D images are generated autoregressively in batches along sampled locations derived from the floorplan. At each step, the diffusion model conditions on previously generated images to produce new images at nearby locations. The global floorplan and attention design in the diffusion model ensures the consistency of the generated images, from which a 3D scene can be reconstructed. Through extensive evaluation on the 3D-FRONT dataset, we demonstrate that HouseCrafter can generate high-quality house-scale 3D scenes. Ablation studies also validate the effectiveness of different design choices. We will release our code and model weights.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5286499-d5f4-4bed-a501-a788ab81b241Cited by top-tier papers3
- From Programs to Poses: Factored Real-World Scene Generation via Learned Program LibrariesJoy Hsu, Emily Jin, Jiajun Wu, Niloy J. MitraNeurIPS 2025 · 6 citations
- Raster2Seq: Polygon Sequence Generation for Floorplan ReconstructionHao Phung, Hadar Averbuch-ElorSIGGRAPH 2026 · 2 citations
- ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image IntermediaryZeqi Gu, Yin Cui, Zhaoshuo Li, Fangyin Wei et al.CVPR 2025
Builds on51
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
Related papers
- RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and GenerationTitas Anciukevicius, Zexiang Xu, Matthew Fisher, Paul Henderson et al.CVPR 2023
- MVDream: Multi-view Diffusion for 3D GenerationYichun Shi, Peng Wang, Jianglong Ye, Long Mai et al.ICLR 2024 · 973 citations
- ConsistDreamer: 3D-Consistent 2D Diffusion for High-Fidelity Scene EditingJun-Kun Chen, Samuel Rota Bulò, Norman Müller, Lorenzo Porzi et al.CVPR 2024
- DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat GenerationChenguo Lin, Panwang Pan, Bangbang Yang, Zeming Li et al.ICLR 2025
- Plan2Scene: Converting Floorplans to 3D ScenesMadhawa Vidanapathirana, Qirui Wu, Yasutaka Furukawa, Angel X. Chang et al.CVPR 2021
