Leveraging BEV Paradigm for Ground-to-Aerial Image Synthesis
Junyan Ye, Jun He, Weijia Li, Zhutao Lv, Yi Lin, Jinhua Yu, Haote Yang, Conghui He
Abstract
Ground-to-aerial image synthesis focuses on generating realistic aerial images from corresponding ground street view images while maintaining consistent content layout, simulating a top-down view. The significant viewpoint difference leads to domain gaps between views, and dense urban scenes limit the visible range of street views, making this cross-view generation task particularly challenging. In this paper, we introduce SkyDiffusion, a novel cross-view generation method for synthesizing aerial images from street view images, utilizing a diffusion model and the Bird's-Eye View (BEV) paradigm. The Curved-BEV method in SkyDiffusion converts street-view images into a BEV perspective, effectively bridging the domain gap, and employs a "multi-to-one" mapping strategy to address occlusion issues in dense urban scenes. Next, Sky-Diffusion designed a BEV-guided diffusion model to generate content-consistent and realistic aerial images. Additionally, we introduce a novel dataset, Ground2Aerial-3, designed for diverse ground-to-aerial image synthesis applications, including disaster scene aerial synthesis, lowaltitude UAV image synthesis, and historical high-resolution satellite image synthesis tasks. Experimental results demonstrate that SkyDiffusion outperforms state-of-the-art methods on cross-view datasets across natural (CVUSA), suburban (CVACT), urban (VIGOR-Chicago), and various application scenarios (G2A-3), achieving realistic and contentconsistent aerial image generation. The code, datasets and more information of this work can be found at https: //opendatalab.github.io/skydiffusion/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a760f33-665d-4e9c-8507-326e07d30c83Cited by top-tier papers7
- Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationSiwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang et al.NeurIPS 2025 · 82 citations
- UrbanFeel:A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human PerspectiveJun He, Yi Lin, Zilong Huang, Jiacong Yin et al.ICLR 2026 · 7 citations
- Geo2: Geometry-Guided Cross-view Geo-Localization and Image SynthesisYancheng Zhang, Xiaohan Zhang, Guangyu Sun, Zonglin Lyu et al.CVPR 2026 · 5 citations
- MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and LayoutsZilong Huang, Jun He, Xiaobin Huang, Ziyi Xiong et al.CVPR 2026 · 5 citations
- SA-Occ: Satellite-Assisted 3D Occupancy Prediction in Real WorldChen Chen, Zhirui Wang, Taowei Sheng, Yi Jiang et al.ICCV 2025 · 1 citation
Builds on24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- MagicDrive: Street View Generation with Diverse 3D Geometry ControlRuiyuan Gao, Kai Chen, Enze Xie, Lanqing Hong et al.ICLR 2024 · 248 citations
Related papers
- Satellite to GroundScape - Large-scale Consistent Ground View Generation from Satellite ViewsNingli Xu, Rongjun QinCVPR 2025
- Bridging the Domain Gap for Ground-to-Aerial Image MatchingKrishna Regmi, Mubarak ShahICCV 2019 · 191 citations
- Sat2Scene: 3D Urban Scene Generation from Satellite Images with DiffusionZuoyue Li, Zhenqiang Li, Zhaopeng Cui, Marc Pollefeys et al.CVPR 2024
- FG^2: Fine-Grained Cross-View Localization by Fine-Grained Feature MatchingZimin Xia, Alexandre AlahiCVPR 2025
- Controllable 3D Outdoor Scene Generation via Scene GraphsYuheng Liu, Xinke Li, Yuning Zhang, Lu Qi et al.ICCV 2025 · 13 citations
