Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
Keyang Lu, Sifan Zhou, Hongbin Xu, Gang Xu, Zhifei Yang, Yikai Wang, Zhen Xiao, Jieyi Long, Ming Li
Abstract
Realistic 3D city generation is fundamental to a wide range of applications, including virtual reality and digital twins. However, most existing methods rely on training a single diffusion model, which limits their ability to generate personalized and boundless city-scale scenes. In this paper, we present Yo'City, a novel agentic framework that enables user-customized and infinitely expandable 3D city generation by leveraging the reasoning and compositional capabilities of off-the-shelf large models. Specifically, Yo'City first conceptualize the city through a top-down planning strategy that defines a hierarchical “City–District–Grid” structure. The Global Planner determines the overall layout and potential functional districts, while the Local Designer further refines each district with detailed grid-level descriptions. Subsequently, the grid-level 3D generation is achieved through a produce–refine–evaluate isometric image synthesis loop, followed by image-to-3D generation. To simulate continuous city evolution, Yo'City further introduces a user-interactive, relationship-guided expansion mechanism, which performs scene graph–based distance- and semantics-aware layout optimization, ensuring spatially coherent city growth. To comprehensively evaluate our method, we construct a diverse benchmark dataset and design six multi-dimensional metrics that assess generation quality from the perspectives of semantics, geometry, texture, and layout. Extensive experiments demonstrate that Yo'City consistently outperforms existing state-of-the-art methods across all evaluation aspects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20071db0-1e18-4793-aa0f-b48da3545aeaBuilds on29
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani et al.NeurIPS 2023 · 462 citations
- ATISS: Autoregressive Transformers for Indoor Scene SynthesisDespoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis et al.NeurIPS 2021 · 293 citations
- Spatial Mental Modeling from Limited ViewsQineng Wang, Baiqiao Yin, Pingyue Zhang, Jianshu Zhang et al.ICLR 2026 · 92 citations
Related papers
- MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and LayoutsZilong Huang, Jun He, Xiaobin Huang, Ziyi Xiong et al.CVPR 2026 · 5 citations
- Sat2City: 3D City Generation from a Single Satellite Image with Cascaded Latent DiffusionTongyan Hua, Lutao Jiang, Ying-Cong Chen, Wufan ZhaoICCV 2025 · 5 citations
- CitySculpt: 3D City Generation from Satellite Imagery with UV DiffusionXingbo Yao, Xuanmin Wang, Hui XiongACM MM 2025
- WorldGrow: Generating Infinite 3D WorldSikuang Li, Chen Yang, Jiemin Fang, Taoran Yi et al.AAAI 2026 · 10 citations
- InfiniCity: Infinite-Scale City SynthesisChieh Hubert Lin, Hsin-Ying Lee, Willi Menapace, Menglei Chai et al.ICCV 2023 · 86 citations
