Leveraging BEV Paradigm for Ground-to-Aerial Image Synthesis
Junyan Ye, Jun He, Weijia Li, Zhutao Lv, Yi Lin, Jinhua Yu, Haote Yang, Conghui He
摘要
Ground-to-aerial image synthesis focuses on generating realistic aerial images from corresponding ground street view images while maintaining consistent content layout, simulating a top-down view. The significant viewpoint difference leads to domain gaps between views, and dense urban scenes limit the visible range of street views, making this cross-view generation task particularly challenging. In this paper, we introduce SkyDiffusion, a novel cross-view generation method for synthesizing aerial images from street view images, utilizing a diffusion model and the Bird's-Eye View (BEV) paradigm. The Curved-BEV method in SkyDiffusion converts street-view images into a BEV perspective, effectively bridging the domain gap, and employs a "multi-to-one" mapping strategy to address occlusion issues in dense urban scenes. Next, Sky-Diffusion designed a BEV-guided diffusion model to generate content-consistent and realistic aerial images. Additionally, we introduce a novel dataset, Ground2Aerial-3, designed for diverse ground-to-aerial image synthesis applications, including disaster scene aerial synthesis, lowaltitude UAV image synthesis, and historical high-resolution satellite image synthesis tasks. Experimental results demonstrate that SkyDiffusion outperforms state-of-the-art methods on cross-view datasets across natural (CVUSA), suburban (CVACT), urban (VIGOR-Chicago), and various application scenarios (G2A-3), achieving realistic and contentconsistent aerial image generation. The code, datasets and more information of this work can be found at https: //opendatalab.github.io/skydiffusion/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationSiwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang 等NeurIPS 2025 · 被引用 82 次
- UrbanFeel:A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human PerspectiveJun He, Yi Lin, Zilong Huang, Jiacong Yin 等ICLR 2026 · 被引用 7 次
- Geo2: Geometry-Guided Cross-view Geo-Localization and Image SynthesisYancheng Zhang, Xiaohan Zhang, Guangyu Sun, Zonglin Lyu 等CVPR 2026 · 被引用 5 次
- MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and LayoutsZilong Huang, Jun He, Xiaobin Huang, Ziyi Xiong 等CVPR 2026 · 被引用 5 次
- SA-Occ: Satellite-Assisted 3D Occupancy Prediction in Real WorldChen Chen, Zhirui Wang, Taowei Sheng, Yi Jiang 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- MagicDrive: Street View Generation with Diverse 3D Geometry ControlRuiyuan Gao, Kai Chen, Enze Xie, Lanqing Hong 等ICLR 2024 · 被引用 248 次
相关 Paper
- Satellite to GroundScape - Large-scale Consistent Ground View Generation from Satellite ViewsNingli Xu, Rongjun QinCVPR 2025
- Bridging the Domain Gap for Ground-to-Aerial Image MatchingKrishna Regmi, Mubarak ShahICCV 2019 · 被引用 191 次
- Sat2Scene: 3D Urban Scene Generation from Satellite Images with DiffusionZuoyue Li, Zhenqiang Li, Zhaopeng Cui, Marc Pollefeys 等CVPR 2024
- FG^2: Fine-Grained Cross-View Localization by Fine-Grained Feature MatchingZimin Xia, Alexandre AlahiCVPR 2025
- Controllable 3D Outdoor Scene Generation via Scene GraphsYuheng Liu, Xinke Li, Yuning Zhang, Lu Qi 等ICCV 2025 · 被引用 13 次
