Sat2Vid: Street-view Panoramic Video Synthesis from a Single Satellite Image
Zuoyue Li, Zhenqiang Li, Zhaopeng Cui, Rongjun Qin, Marc Pollefeys, Martin R. Oswald
Abstract
We present a novel method for synthesizing both temporally and geometrically consistent street-view panoramic video from a single satellite image and camera trajectory. Existing cross-view synthesis approaches focus on images, while video synthesis in such a case has not yet received enough attention. For geometrical and temporal consistency, our approach explicitly creates a 3D point cloud representation of the scene and maintains dense 3D-2D correspondences across frames that reflect the geometric scene configuration inferred from the satellite view. As for synthesis in the 3D space, we implement a cascaded network architecture with two hourglass modules to generate point-wise coarse and fine features from semantics and per-class latent vectors, followed by projection to frames and an up-sampling module to obtain the final realistic video. By leveraging computed correspondences, the produced street-view video frames adhere to the 3D geometric scene structure and maintain temporal consistency. Qualitative and quantitative experiments demonstrate superior results compared to other state-of-the-art synthesis approaches that either lack temporal consistency or realistic appearance. To the best of our knowledge, our work is the first one to synthesize cross-view images to videos..
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d7d3fd5-4a99-4c6a-b891-8e8dcbd8458dCited by top-tier papers9
- Boosting 3-DoF Ground-to-Satellite Camera Localization Accuracy via Geometry-Guided Cross-View TransformerYujiao Shi, Fei Wu, Akhil Perincherry, Ankit Vora et al.ICCV 2023 · 60 citations
- Sat2Density: Faithful Density Learning from Satellite-Ground Image PairsMing Qian, Jincheng Xiong, Gui-Song Xia, Nan XueICCV 2023 · 29 citations
- Sat2City: 3D City Generation from a Single Satellite Image with Cascaded Latent DiffusionTongyan Hua, Lutao Jiang, Ying-Cong Chen, Wufan ZhaoICCV 2025 · 5 citations
- Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite ImageMing Qian, Zimin Xia, Changkun Liu, Shuailei Ma et al.ICLR 2026 · 5 citations
- SatDreamer360: Multiview-Consistent Generation of Ground-Level Scenes from Satellite ImageryXianghui Ze, Beiyi Zhu, Zhenbo Song, Jianfeng Lu et al.ICLR 2026 · 1 citation
Builds on6
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Bridging the Domain Gap for Ground-to-Aerial Image MatchingKrishna Regmi, Mubarak ShahICCV 2019 · 191 citations
- Geometry-Aware Satellite-to-Ground Image Synthesis for Urban AreasXiaohu Lu, Zuoyue Li, Zhaopeng Cui, Martin R. Oswald et al.CVPR 2020
- RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point CloudsQingyong Hu, Bo Yang, Linhai Xie, Stefano Rosa et al.CVPR 2020
- SynSin: End-to-End View Synthesis From a Single ImageOlivia Wiles, Georgia Gkioxari, Richard Szeliski, Justin JohnsonCVPR 2020
Related papers
- Coming Down to Earth: Satellite-to-Street View Synthesis for Geo-LocalizationAysim Toker, Qunjie Zhou, Maxim Maximov, Laura Leal-TaixéCVPR 2021
- Satellite to GroundScape - Large-scale Consistent Ground View Generation from Satellite ViewsNingli Xu, Rongjun QinCVPR 2025
- Dynamic View Synthesis with Spatio-Temporal Feature Warping from Sparse ViewsDeqi Li, Shi-Sheng Huang, Tianyu Shen, Hua HuangACM MM 2023 · 6 citations
- Look Beyond: Two-Stage Scene View Generation via Panorama and Video DiffusionXueyang Kang, Zhengkang Xiang, Zezheng Zhang, Kourosh KhoshelhamACM MM 2025
- Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single ImageXuanchi Ren, Xiaolong WangCVPR 2022 · 42 citations
