UAVScenes: A Multi-Modal Dataset for UAVs
Sijie Wang, Siqi Li, Yawei Zhang, Shangshu Yu, Shenghai Yuan, Rui She, Quanjiang Guo, Jinxuan Zheng, Ong Kang Howe, Leonrich Chandra, Shrivarshann Srijeyan, Aditya Sivadas
Abstract
Multi-modal perception is essential for unmanned aerial vehicle (UAV) operations, as it enables a comprehensive understanding of the UAVs' surrounding environment. However, most existing multi-modal UAV datasets are primarily biased toward localization and 3D reconstruction tasks, or only support map-level semantic segmentation due to the lack of frame-wise annotations for both camera images and LiDAR point clouds. This limitation prevents them from being used for high-level scene understanding tasks. To address this gap and advance multi-modal UAV perception, we introduce UAVScenes, a large-scale dataset designed to benchmark various tasks across both 2D and 3D modalities. Our benchmark dataset is built upon the well-calibrated multi-modal UAV dataset MARS-LVIG, originally developed only for simultaneous localization and mapping (SLAM). We enhance this dataset by providing manually labeled semantic annotations for both frame-wise images and LiDAR point clouds, along with accurate 6-degree-of-freedom (6-DoF) poses. These additions enable a wide range of UAV perception tasks, including segmentation, depth estimation, 6-DoF localization, place recognition, and novel view synthesis (NVS). Our dataset is available at https://github.com/sijieaaa/ UAVScenes
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f891b8c-0772-43cf-aba1-4d07e8aabe47Cited by top-tier papers4
- AirSim360: A Panoramic Simulation Platform within Drone ViewXian Ge, Yuling Pan, Yuhang Zhang, Xiang Li et al.CVPR 2026 · 17 citations
- Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV NavigationLingfeng Zhang, Yuchen Zhang, Hongsheng Li, Haoxiang Fu et al.CVPR 2026 · 15 citations
- PiLoT: Neural Pixel-to-3D Registration for UAV-based Ego and Target Geo-localizationXiaoya Cheng, Long Wang, Yan Liu, Xinyi Liu et al.CVPR 2026 · 6 citations
- AeroDGS: Physically Consistent Dynamic Gaussian Splatting for Single-Sequence Aerial 4D ReconstructionHanyang Liu, Rongjun QinCVPR 2026 · 2 citations
Builds on26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
Related papers
- Human-centric Scene Understanding for 3D Large-scale ScenariosYiteng Xu, Peishan Cong, Yichen Yao, Runnan Chen et al.ICCV 2023 · 34 citations
- V2U4Real: A Real-world Large-scale Dataset for Vehicle-to-UAV Cooperative PerceptionWeijia Li, Haoen Xiang, Tianxu Wang, Shuaibing Wu et al.CVPR 2026 · 4 citations
- U2UData: A Large-scale Cooperative Perception Dataset for Swarm UAVs Autonomous FlightTongtong Feng, Xin Wang, Feilin Han, Leping Zhang et al.ACM MM 2024 · 19 citations
- PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning SegmentationShuyan Ke, Yifan Mei, Changli Wu, Yonghan Zheng et al.CVPR 2026 · 3 citations
- City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete LearningPenglei Sun, Yaoxian Song, Xiangru Zhu, Xiang Liu et al.ACM MM 2025 · 2 citations
