UAVScenes: A Multi-Modal Dataset for UAVs
Sijie Wang, Siqi Li, Yawei Zhang, Shangshu Yu, Shenghai Yuan, Rui She, Quanjiang Guo, Jinxuan Zheng, Ong Kang Howe, Leonrich Chandra, Shrivarshann Srijeyan, Aditya Sivadas
摘要
Multi-modal perception is essential for unmanned aerial vehicle (UAV) operations, as it enables a comprehensive understanding of the UAVs' surrounding environment. However, most existing multi-modal UAV datasets are primarily biased toward localization and 3D reconstruction tasks, or only support map-level semantic segmentation due to the lack of frame-wise annotations for both camera images and LiDAR point clouds. This limitation prevents them from being used for high-level scene understanding tasks. To address this gap and advance multi-modal UAV perception, we introduce UAVScenes, a large-scale dataset designed to benchmark various tasks across both 2D and 3D modalities. Our benchmark dataset is built upon the well-calibrated multi-modal UAV dataset MARS-LVIG, originally developed only for simultaneous localization and mapping (SLAM). We enhance this dataset by providing manually labeled semantic annotations for both frame-wise images and LiDAR point clouds, along with accurate 6-degree-of-freedom (6-DoF) poses. These additions enable a wide range of UAV perception tasks, including segmentation, depth estimation, 6-DoF localization, place recognition, and novel view synthesis (NVS). Our dataset is available at https://github.com/sijieaaa/ UAVScenes
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- AirSim360: A Panoramic Simulation Platform within Drone ViewXian Ge, Yuling Pan, Yuhang Zhang, Xiang Li 等CVPR 2026 · 被引用 17 次
- Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV NavigationLingfeng Zhang, Yuchen Zhang, Hongsheng Li, Haoxiang Fu 等CVPR 2026 · 被引用 15 次
- PiLoT: Neural Pixel-to-3D Registration for UAV-based Ego and Target Geo-localizationXiaoya Cheng, Long Wang, Yan Liu, Xinyi Liu 等CVPR 2026 · 被引用 6 次
- AeroDGS: Physically Consistent Dynamic Gaussian Splatting for Single-Sequence Aerial 4D ReconstructionHanyang Liu, Rongjun QinCVPR 2026 · 被引用 2 次
它引用的顶会 Paper26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
相关 Paper
- Human-centric Scene Understanding for 3D Large-scale ScenariosYiteng Xu, Peishan Cong, Yichen Yao, Runnan Chen 等ICCV 2023 · 被引用 34 次
- V2U4Real: A Real-world Large-scale Dataset for Vehicle-to-UAV Cooperative PerceptionWeijia Li, Haoen Xiang, Tianxu Wang, Shuaibing Wu 等CVPR 2026 · 被引用 4 次
- U2UData: A Large-scale Cooperative Perception Dataset for Swarm UAVs Autonomous FlightTongtong Feng, Xin Wang, Feilin Han, Leping Zhang 等ACM MM 2024 · 被引用 19 次
- PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning SegmentationShuyan Ke, Yifan Mei, Changli Wu, Yonghan Zheng 等CVPR 2026 · 被引用 3 次
- City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete LearningPenglei Sun, Yaoxian Song, Xiangru Zhu, Xiang Liu 等ACM MM 2025 · 被引用 2 次
