Sky2Ground: A Benchmark for Site Modeling under Varying Altitude
Zengyan Wang, Sirshapan Mitra, Rajat Modi, Hui Xian Grace Lim, Yogesh Rawat
摘要
We introduce Sky2Ground, a three-view dataset designed for varying altitude camera localization, correspondence learning, and reconstruction. The dataset combines structured synthetic imagery with real, in-the-wild images, providing both controlled multi-view geometry and realistic scene noise. Each of the 51 sites contains thousands of satellite, aerial, and ground images spanning wide altitude ranges and nearly orthogonal viewing angles, enabling rigorous evaluation across global-to-local contexts. We benchmark state of the art pose estimation models, including MASt3R, DUSt3R, Map Anything, and VGGT, and observe that the use of satellite imagery often degrades performance, highlighting the challenges under large altitude variations. We also examine reconstruction methods, highlighting the challenges introduced by sparse geometric overlap, varying perspectives, and the use of real imagery, which often introduces noise and reduces rendering quality. To address some of these challenges, we propose SkyNet, a model which enhances cross-view consistency when incorporating satellite imagery with a curriculum-based training strategy to progressively incorporate more satellite views. SkyNet significantly strengthens multi-view alignment and outperforms existing methods by 9.6% on RRA@5 and 18.1% on RTA@5 in terms of absolute performance. Sky2Ground and SkyNet together establish a comprehensive testbed and baseline for advancing large-scale, multi-altitude 3D perception and generalizable camera localization. Code and models will be released publicly for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- MatrixCity: A Large-scale City Dataset for City-scale Neural Rendering and BeyondYixuan Li, Lihan Jiang, Linning Xu, Yuanbo Xiangli 等ICCV 2023 · 被引用 185 次
- CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical FlowPhilippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon 等ICCV 2023 · 被引用 181 次
- Beyond Geo-localization: Fine-grained Orientation of Street-view Images by Cross-view Matching with Satellite ImageryWenmiao Hu, Yichen Zhang, Yuxuan Liang, Yifang Yin 等ACM MM 2022 · 被引用 31 次
相关 Paper
- SkyNet: Multi-Drone Cooperation for Real-Time Person Identification and LocalizationJunkun Peng, Qing Li, Yuanzheng Tan, Dan Zhao 等INFOCOM 2023 · 被引用 10 次
- Learning Dense Flow Field for Highly-accurate Cross-view Camera LocalizationZhenbo Song, Xianghui Ze, Jianfeng Lu, Yujiao ShiNeurIPS 2023 · 被引用 37 次
- Cross-View Visual Geo-Localization for Outdoor Augmented RealityNiluthpol Chowdhury Mithun, Kshitij Minhas, Han-Pang Chiu, Taragay Oskiper 等IEEE VR 2023 · 被引用 22 次
- MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsZhenggang Tang, Yuchen Fan, Dilin Wang, Hongyu Xu 等CVPR 2025
- Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced ImagesMatias Turkulainen, Akshay Krishnan, Filippo Aleotti, Mohamed Sayed 等CVPR 2026
