Architect: Generating Vivid and Interactive 3D Scenes with Hierarchical 2D Inpainting
Yian Wang, Xiaowen Qiu, Jiageng Liu, Zhehuan Chen, Jiting Cai, Yufei Wang, Tsun-Hsuan Johnson Wang, Zhou Xian, Chuang Gan
Abstract
Creating large-scale interactive 3D environments is essential for the development of Robotics and Embodied AI research. Current methods, including manual design, procedural generation, diffusion-based scene generation, and large language model (LLM) guided scene design, are hindered by limitations such as excessive human effort, reliance on predefined rules or training datasets, and limited 3D spatial reasoning ability. Since pre-trained 2D image generative models better capture scene and object configuration than LLMs, we address these challenges by introducing Architect, a generative framework that creates complex and realistic 3D embodied environments leveraging diffusion-based 2D image inpainting. In detail, we utilize foundation visual perception models to obtain each generated object from the image and leverage pre-trained depth estimation models to lift the generated 2D image to 3D space. Our pipeline is further extended to a hierarchical and iterative inpainting process to continuously generate placement of large furniture and small objects to enrich the scene. This iterative structure brings the flexibility for our method to generate or refine scenes from various starting points, such as text, floor plans, or pre-arranged environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8874f4f9-2bc4-4ef9-ab09-a0d0352bcd35Cited by top-tier papers14
- Scenethesis: A Language and Vision Agentic Framework for 3D Scene GenerationLu Ling, Chen-Hsuan Lin, Tsung-Yi Lin, Yifan Ding et al.ICLR 2026 · 74 citations
- SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective AgentYandan Yang, Baoxiong Jia, Shujie Zhang, Siyuan HuangNeurIPS 2025 · 65 citations
- SAGE: Scalable Agentic 3D Scene Generation for Embodied AIHongchi Xia, Xuan Li, Zhaoshuo Li, Qianli Ma et al.CVPR 2026 · 50 citations
- MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial ReasoningJinkun Hao, Naifu Liang, Zhen Luo, Xudong Xu et al.NeurIPS 2025 · 21 citations
- HoloScene: Simulation-Ready Interactive 3D Worlds from a Single VideoHongchi Xia, Chih-Hao Lin, Hao-Yu Hsu, Quentin Leboutet et al.NeurIPS 2025 · 18 citations
Builds on23
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
Related papers
- WorldGen: From Text to Traversable and Interactive 3D WorldsDilin Wang, Hyunyoung Jung, Tom Monnier, Kihyuk Sohn et al.CVPR 2026 · 24 citations
- Holodeck: Language Guided Generation of 3D Embodied AI EnvironmentsYue Yang, Fan-Yun Sun, Luca Weihs, Eli VanderBilt et al.CVPR 2024 · 47 citations
- 3D-VLA: A 3D Vision-Language-Action Generative World ModelHaoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang et al.ICML 2024 · 303 citations
- NeuralField-LDM: Scene Generation with Hierarchical Latent Diffusion ModelsSeung Wook Kim, Bradley Brown, Kangxue Yin, Karsten Kreis et al.CVPR 2023
- SPREAD: Spatial-Physical REasoning via geometry Aware DiffusionMinzhang Li, Kuixiang Shao, Xuebing Li, Yuyang Jiao et al.CVPR 2026
