SatDreamer360: Multiview-Consistent Generation of Ground-Level Scenes from Satellite Imagery
Xianghui Ze, Beiyi Zhu, Zhenbo Song, Jianfeng Lu, Yujiao Shi
Abstract
Generating multiview-consistent 360 • ground-level scenes from satellite imagery is a challenging task with broad applications in simulation, autonomous navigation, and digital twin cities. Existing approaches primarily focus on synthesizing individual ground-view panoramas, often relying on auxiliary inputs like height maps or handcrafted projections, and struggle to produce multiview consistent sequences. In this paper, we propose SatDreamer360, a framework that generates geometrically consistent multi-view ground-level panoramas from a single satellite image, given a predefined pose trajectory. To address the large viewpoint discrepancy between ground and satellite images, we adopt a triplane representation to encode scene features and design a ray-based pixel attention mechanism that retrieves view-specific features from the triplane. To maintain multi-frame consistency, we introduce a panoramic epipolar-constrained attention module that aligns features across frames based on known relative poses. To support the evaluation, we introduce VIGOR++, a large-scale dataset for generating multi-view ground panoramas from a satellite image, by augmenting the original VIGOR dataset with more ground-view images and their pose annotations. Experiments show that SatDreamer360 outperforms existing methods in both satellite-to-ground alignment and multiview consistency. INTRODUCTION Generating ground-level scenes from satellite imagery has attracted significant attention due to the broad coverage and low acquisition cost of satellite images. This task shows promising applications in autonomous driving (Villalonga Pineda ( 2021 ); Lu et al. (2024)), 3D reconstruction (Liu et al. (2024); Yan et al. (2024)) and data augmentation (Yang et al. (2023); Gao et al. (2023)) for downstream tasks. Many existing works (Li et al. (2024a); Lin et al. (2024); Xu & Qin (2024); Ze et al. (2025)) focus on generating individual ground images from satellite views, leaving the continuity of multi-ground views largely unaddressed. In this paper, we aim to synthesize multiple ground-view images from a single satellite image, controlled by a predefined trajectory. This introduces new challenges in maintaining both geometric consistency with the top-down satellite image and multiview coherence across the sequence of generated frames. Early approaches (Isola et al. (2017a); Regmi & Borji (2018); Shi et al. (2022); Lu et al. (2020); Qian et al. ( 2023 )) formulate cross-view synthesis as a one-to-one mapping problem, often implemented with Conditional Generative Adversarial Networks (cGANs). These methods focus on aligning representations at pixel or perceptual level. However, the extreme viewpoint disparity between top-down satellite views and street-level images leads to limited field-of-view overlap. Satellite images inherently miss key elements such as building facades, tree trunks, and other occluded details, making the ground view generation task highly under-constrained and naturally one-to-many. Recent advances leverage latent diffusion models (LDMs) (Rombach et al. (2022)) to better handle this uncertainty (Li et al. (2024a); Lin et al. (2024); Deng et al. (2024); Xu & Qin (2024); Ze et al. ( 2025 )). These methods introduce probabilistic modeling to produce diverse and high-fidelity ground
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0d6ec2a-700c-4e00-bb3f-27f8d4ffec6eCited by top-tier papers2
- BevSplat: Resolving Height Ambiguity via Feature-Based Gaussian Primitives for Weakly-Supervised Cross-View LocalizationQiwei Wang, Shaoxun Wu, Yujiao ShiNeurIPS 2025 · 10 citations
- Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced ImagesMatias Turkulainen, Akshay Krishnan, Filippo Aleotti, Mohamed Sayed et al.CVPR 2026
Builds on31
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Satellite to GroundScape - Large-scale Consistent Ground View Generation from Satellite ViewsNingli Xu, Rongjun QinCVPR 2025
- Geometry-Aware Satellite-to-Ground Image Synthesis for Urban AreasXiaohu Lu, Zuoyue Li, Zhaopeng Cui, Martin R. Oswald et al.CVPR 2020
- Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite ImageMing Qian, Zimin Xia, Changkun Liu, Shuailei Ma et al.ICLR 2026 · 5 citations
- Sat2Vid: Street-view Panoramic Video Synthesis from a Single Satellite ImageZuoyue Li, Zhenqiang Li, Zhaopeng Cui, Rongjun Qin et al.ICCV 2021 · 26 citations
- VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose EstimationJuhye Park, Wooju Lee, Dasol Hong, Changki Sung et al.CVPR 2026
