SpaceNet MVOI: A Multi-View Overhead Imagery Dataset
Nicholas Weir, David Lindenbaum, Alexei Bastidas, Adam Van Etten, Varun Kumar Vijay, Sean McPherson, Jacob Shermeyer, Hanlin Tang
摘要
Detection and segmentation of objects in overheard imagery is a challenging task. The variable density, random orientation, small size, and instance-to-instance heterogeneity of objects in overhead imagery calls for approaches distinct from existing models designed for natural scene datasets. Though new overhead imagery datasets are being developed, they almost universally comprise a single view taken from directly overhead ("at nadir"), failing to address a critical variable: look angle. By contrast, views vary in real-world overhead imagery, particularly in dynamic scenarios such as natural disasters where first looks are often over 40 degrees off-nadir. This represents an important challenge to computer vision methods, as changing view angle adds distortions, alters resolution, and changes lighting. At present, the impact of these perturbations for algorithmic detection and segmentation of objects is untested. To address this problem, we present an open source Multi-View Overhead Imagery dataset, termed SpaceNet MVOI, with 27 unique looks from a broad range of viewing angles (-32.5 degrees to 54.0 degrees). Each of these images cover the same 665 square km geographic extent and are annotated with 126,747 building footprint labels, enabling direct assessment of the impact of viewpoint perturbation on model performance. We benchmark multiple leading segmentation and object detection models on: (1) building detection, (2) generalization to unseen viewing angles and resolutions, and (3) sensitivity of building footprint extraction to changes in resolution. We find that state of the art segmentation and object detection models struggle to identify buildings in off-nadir imagery and generalize poorly to unseen views, presenting an important benchmark to explore the broadly relevant challenge of detecting small, heterogeneous target objects in visually dynamic contexts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- DynamicEarthNet: Daily Multi-Spectral Satellite Dataset for Semantic Change SegmentationAysim Toker, Lukas Kondmann, Mark Weber, Marvin Eisenberger 等CVPR 2022 · 被引用 108 次
- 3D Building Reconstruction from Monocular Remote Sensing ImagesWeijia Li, Lingxuan Meng, Jinwang Wang, Conghui He 等ICCV 2021 · 被引用 46 次
- CityDreamer: Compositional Generative Model of Unbounded 3D CitiesHaozhe Xie, Zhaoxi Chen, Fangzhou Hong, Ziwei LiuCVPR 2024 · 被引用 35 次
- HyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing ImageryJingtao Li, Yingyi Liu, Xinyu Wang, Yunning Peng 等CVPR 2025
- The Multi-Temporal Urban Development SpaceNet DatasetAdam Van Etten, Daniel Hogan, Jesus Martinez-Manso, Jacob Shermeyer 等CVPR 2021
相关 Paper
- MOR-UAV: A Benchmark Dataset and Baselines for Moving Object Recognition in UAV VideosMurari Mandal, Lav Kush Kumar, Santosh Kumar VipparthiACM MM 2020 · 被引用 58 次
- Sky2Ground: A Benchmark for Site Modeling under Varying AltitudeZengyan Wang, Sirshapan Mitra, Rajat Modi, Hui Xian Grace Lim 等CVPR 2026
- Augmenting Depth Estimation with Geospatial ContextScott Workman, Hunter BlantonICCV 2021 · 被引用 6 次
- OpenStreetView-5M: The Many Roads to Global Visual GeolocationGuillaume Astruc, Nicolas Dufour, Ioannis Siglidis, Constantin Aronssohn 等CVPR 2024
- Boundary-Aware 3D Building Reconstruction From a Single Overhead ImageJisan Mahmud, True Price, Akash Bapat, Jan-Michael FrahmCVPR 2020
