Sky2Ground: A Benchmark for Site Modeling under Varying Altitude
Zengyan Wang, Sirshapan Mitra, Rajat Modi, Hui Xian Grace Lim, Yogesh Rawat
Abstract
We introduce Sky2Ground, a three-view dataset designed for varying altitude camera localization, correspondence learning, and reconstruction. The dataset combines structured synthetic imagery with real, in-the-wild images, providing both controlled multi-view geometry and realistic scene noise. Each of the 51 sites contains thousands of satellite, aerial, and ground images spanning wide altitude ranges and nearly orthogonal viewing angles, enabling rigorous evaluation across global-to-local contexts. We benchmark state of the art pose estimation models, including MASt3R, DUSt3R, Map Anything, and VGGT, and observe that the use of satellite imagery often degrades performance, highlighting the challenges under large altitude variations. We also examine reconstruction methods, highlighting the challenges introduced by sparse geometric overlap, varying perspectives, and the use of real imagery, which often introduces noise and reduces rendering quality. To address some of these challenges, we propose SkyNet, a model which enhances cross-view consistency when incorporating satellite imagery with a curriculum-based training strategy to progressively incorporate more satellite views. SkyNet significantly strengthens multi-view alignment and outperforms existing methods by 9.6% on RRA@5 and 18.1% on RTA@5 in terms of absolute performance. Sky2Ground and SkyNet together establish a comprehensive testbed and baseline for advancing large-scale, multi-altitude 3D perception and generalizable camera localization. Code and models will be released publicly for future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9d96378-6b3c-4358-ae36-e28088a8ddf8Builds on14
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- MatrixCity: A Large-scale City Dataset for City-scale Neural Rendering and BeyondYixuan Li, Lihan Jiang, Linning Xu, Yuanbo Xiangli et al.ICCV 2023 · 185 citations
- CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical FlowPhilippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon et al.ICCV 2023 · 181 citations
- Beyond Geo-localization: Fine-grained Orientation of Street-view Images by Cross-view Matching with Satellite ImageryWenmiao Hu, Yichen Zhang, Yuxuan Liang, Yifang Yin et al.ACM MM 2022 · 31 citations
Related papers
- SkyNet: Multi-Drone Cooperation for Real-Time Person Identification and LocalizationJunkun Peng, Qing Li, Yuanzheng Tan, Dan Zhao et al.INFOCOM 2023 · 10 citations
- Learning Dense Flow Field for Highly-accurate Cross-view Camera LocalizationZhenbo Song, Xianghui Ze, Jianfeng Lu, Yujiao ShiNeurIPS 2023 · 37 citations
- Cross-View Visual Geo-Localization for Outdoor Augmented RealityNiluthpol Chowdhury Mithun, Kshitij Minhas, Han-Pang Chiu, Taragay Oskiper et al.IEEE VR 2023 · 22 citations
- MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsZhenggang Tang, Yuchen Fan, Dilin Wang, Hongyu Xu et al.CVPR 2025
- Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced ImagesMatias Turkulainen, Akshay Krishnan, Filippo Aleotti, Mohamed Sayed et al.CVPR 2026
