Sat2Vid: Street-view Panoramic Video Synthesis from a Single Satellite Image
Zuoyue Li, Zhenqiang Li, Zhaopeng Cui, Rongjun Qin, Marc Pollefeys, Martin R. Oswald
摘要
We present a novel method for synthesizing both temporally and geometrically consistent street-view panoramic video from a single satellite image and camera trajectory. Existing cross-view synthesis approaches focus on images, while video synthesis in such a case has not yet received enough attention. For geometrical and temporal consistency, our approach explicitly creates a 3D point cloud representation of the scene and maintains dense 3D-2D correspondences across frames that reflect the geometric scene configuration inferred from the satellite view. As for synthesis in the 3D space, we implement a cascaded network architecture with two hourglass modules to generate point-wise coarse and fine features from semantics and per-class latent vectors, followed by projection to frames and an up-sampling module to obtain the final realistic video. By leveraging computed correspondences, the produced street-view video frames adhere to the 3D geometric scene structure and maintain temporal consistency. Qualitative and quantitative experiments demonstrate superior results compared to other state-of-the-art synthesis approaches that either lack temporal consistency or realistic appearance. To the best of our knowledge, our work is the first one to synthesize cross-view images to videos..
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Boosting 3-DoF Ground-to-Satellite Camera Localization Accuracy via Geometry-Guided Cross-View TransformerYujiao Shi, Fei Wu, Akhil Perincherry, Ankit Vora 等ICCV 2023 · 被引用 60 次
- Sat2Density: Faithful Density Learning from Satellite-Ground Image PairsMing Qian, Jincheng Xiong, Gui-Song Xia, Nan XueICCV 2023 · 被引用 29 次
- Sat2City: 3D City Generation from a Single Satellite Image with Cascaded Latent DiffusionTongyan Hua, Lutao Jiang, Ying-Cong Chen, Wufan ZhaoICCV 2025 · 被引用 5 次
- Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite ImageMing Qian, Zimin Xia, Changkun Liu, Shuailei Ma 等ICLR 2026 · 被引用 5 次
- SatDreamer360: Multiview-Consistent Generation of Ground-Level Scenes from Satellite ImageryXianghui Ze, Beiyi Zhu, Zhenbo Song, Jianfeng Lu 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper6
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 被引用 840 次
- Bridging the Domain Gap for Ground-to-Aerial Image MatchingKrishna Regmi, Mubarak ShahICCV 2019 · 被引用 191 次
- Geometry-Aware Satellite-to-Ground Image Synthesis for Urban AreasXiaohu Lu, Zuoyue Li, Zhaopeng Cui, Martin R. Oswald 等CVPR 2020
- RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point CloudsQingyong Hu, Bo Yang, Linhai Xie, Stefano Rosa 等CVPR 2020
- SynSin: End-to-End View Synthesis From a Single ImageOlivia Wiles, Georgia Gkioxari, Richard Szeliski, Justin JohnsonCVPR 2020
相关 Paper
- Coming Down to Earth: Satellite-to-Street View Synthesis for Geo-LocalizationAysim Toker, Qunjie Zhou, Maxim Maximov, Laura Leal-TaixéCVPR 2021
- Satellite to GroundScape - Large-scale Consistent Ground View Generation from Satellite ViewsNingli Xu, Rongjun QinCVPR 2025
- Dynamic View Synthesis with Spatio-Temporal Feature Warping from Sparse ViewsDeqi Li, Shi-Sheng Huang, Tianyu Shen, Hua HuangACM MM 2023 · 被引用 6 次
- Look Beyond: Two-Stage Scene View Generation via Panorama and Video DiffusionXueyang Kang, Zhengkang Xiang, Zezheng Zhang, Kourosh KhoshelhamACM MM 2025
- Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single ImageXuanchi Ren, Xiaolong WangCVPR 2022 · 被引用 42 次
