SpaceNet MVOI: A Multi-View Overhead Imagery Dataset
Nicholas Weir, David Lindenbaum, Alexei Bastidas, Adam Van Etten, Varun Kumar Vijay, Sean McPherson, Jacob Shermeyer, Hanlin Tang
Abstract
Detection and segmentation of objects in overheard imagery is a challenging task. The variable density, random orientation, small size, and instance-to-instance heterogeneity of objects in overhead imagery calls for approaches distinct from existing models designed for natural scene datasets. Though new overhead imagery datasets are being developed, they almost universally comprise a single view taken from directly overhead ("at nadir"), failing to address a critical variable: look angle. By contrast, views vary in real-world overhead imagery, particularly in dynamic scenarios such as natural disasters where first looks are often over 40 degrees off-nadir. This represents an important challenge to computer vision methods, as changing view angle adds distortions, alters resolution, and changes lighting. At present, the impact of these perturbations for algorithmic detection and segmentation of objects is untested. To address this problem, we present an open source Multi-View Overhead Imagery dataset, termed SpaceNet MVOI, with 27 unique looks from a broad range of viewing angles (-32.5 degrees to 54.0 degrees). Each of these images cover the same 665 square km geographic extent and are annotated with 126,747 building footprint labels, enabling direct assessment of the impact of viewpoint perturbation on model performance. We benchmark multiple leading segmentation and object detection models on: (1) building detection, (2) generalization to unseen viewing angles and resolutions, and (3) sensitivity of building footprint extraction to changes in resolution. We find that state of the art segmentation and object detection models struggle to identify buildings in off-nadir imagery and generalize poorly to unseen views, presenting an important benchmark to explore the broadly relevant challenge of detecting small, heterogeneous target objects in visually dynamic contexts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21941fb9-fb1f-42be-ba45-a192d13feafeCited by top-tier papers8
- DynamicEarthNet: Daily Multi-Spectral Satellite Dataset for Semantic Change SegmentationAysim Toker, Lukas Kondmann, Mark Weber, Marvin Eisenberger et al.CVPR 2022 · 108 citations
- 3D Building Reconstruction from Monocular Remote Sensing ImagesWeijia Li, Lingxuan Meng, Jinwang Wang, Conghui He et al.ICCV 2021 · 46 citations
- CityDreamer: Compositional Generative Model of Unbounded 3D CitiesHaozhe Xie, Zhaoxi Chen, Fangzhou Hong, Ziwei LiuCVPR 2024 · 35 citations
- HyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing ImageryJingtao Li, Yingyi Liu, Xinyu Wang, Yunning Peng et al.CVPR 2025
- The Multi-Temporal Urban Development SpaceNet DatasetAdam Van Etten, Daniel Hogan, Jesus Martinez-Manso, Jacob Shermeyer et al.CVPR 2021
Related papers
- MOR-UAV: A Benchmark Dataset and Baselines for Moving Object Recognition in UAV VideosMurari Mandal, Lav Kush Kumar, Santosh Kumar VipparthiACM MM 2020 · 58 citations
- Sky2Ground: A Benchmark for Site Modeling under Varying AltitudeZengyan Wang, Sirshapan Mitra, Rajat Modi, Hui Xian Grace Lim et al.CVPR 2026
- Augmenting Depth Estimation with Geospatial ContextScott Workman, Hunter BlantonICCV 2021 · 6 citations
- OpenStreetView-5M: The Many Roads to Global Visual GeolocationGuillaume Astruc, Nicolas Dufour, Ioannis Siglidis, Constantin Aronssohn et al.CVPR 2024
- Boundary-Aware 3D Building Reconstruction From a Single Overhead ImageJisan Mahmud, True Price, Akash Bapat, Jan-Michael FrahmCVPR 2020
