Sat2City: 3D City Generation from a Single Satellite Image with Cascaded Latent Diffusion
Tongyan Hua, Lutao Jiang, Ying-Cong Chen, Wufan Zhao
摘要
Recent advancements in generative models have enabled 3D urban scene generation from satellite imagery, unlocking promising applications in gaming, digital twins, and beyond. However, most existing methods rely heavily on neural rendering techniques, which hinder their ability to produce detailed 3D structures on a broader scale, largely due to the inherent structural ambiguity derived from relatively limited 2D observations. To address this challenge, we propose Sat2City, a novel framework that synergizes the representational capacity of sparse voxel grids with latent diffusion models, tailored specifically for our novel 3D city dataset. Our approach is enabled by three key components: (1) A cascaded latent diffusion framework that progressively recovers 3D city structures from satellite imagery, (2) a Re-Hash operation at its Variational Autoencoder (VAE) bottleneck to compute multi-scale feature grids for stable appearance optimization and (3) an inverse sampling strategy enabling implicit supervision for smooth appearance transitioning. To overcome the challenge of collecting real-world city-scale 3D models with high-quality geometry and appearance, we introduce a dataset of synthesized large-scale 3D cities paired with satellite-view height maps. Validated on this dataset, our framework generates detailed 3D structures from a single satellite image, achieving superior fidelity compared to existing city generation models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic ExpansionKeyang Lu, Sifan Zhou, Hongbin Xu, Gang Xu 等CVPR 2026 · 被引用 9 次
- Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite ImageMing Qian, Zimin Xia, Changkun Liu, Shuailei Ma 等ICLR 2026 · 被引用 5 次
- BuildAnyPoint: 3D Building Structured Abstraction from Diverse Point CloudsTongyan Hua, Haoran Gong, Yuan Liu, Di Wang 等CVPR 2026 · 被引用 1 次
- DiMeR: Disentangled Mesh Reconstruction Model with Normal-only Geometry TrainingLutao Jiang, Jiantao Lin, Kanghao Chen, Wenhang Ge 等ICLR 2026
它引用的顶会 Paper45
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields ReconstructionCheng Sun, Min Sun, Hwann-Tzong ChenCVPR 2022 · 被引用 859 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
相关 Paper
- Sat2Scene: 3D Urban Scene Generation from Satellite Images with DiffusionZuoyue Li, Zhenqiang Li, Zhaopeng Cui, Marc Pollefeys 等CVPR 2024
- CitySculpt: 3D City Generation from Satellite Imagery with UV DiffusionXingbo Yao, Xuanmin Wang, Hui XiongACM MM 2025
- ScenDi: 3D-to-2D Scene Diffusion Cascades for Urban GenerationHanlei Guo, Jiahao Shao, Xinya Chen, Xiyang Tan 等CVPR 2026 · 被引用 1 次
- Large Scene Generation with Cube-Absorb Discrete DiffusionQianjiang Hu Wei Hu, Wei HuICCV 2025 · 被引用 3 次
- Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes ModelingZhihao Li, Yufei Wang, Heliang Zheng, Yihao Luo 等NeurIPS 2025 · 被引用 92 次
