Multiview Aerial Visual Recognition (MAVREC): Can Multi-View Improve Aerial Visual Perception?
Aritra Dutta, Srijan Das, Jacob Nielsen, Rajatsubhra Chakraborty, Mubarak Shah
Abstract
Figure 1. Illustration of the geography-aware model using our proposed MAVREC dataset (green box) collected in the rural and urban European landscape vs. the conventional aerial object detector (blue box) pretrained only on aerial images from VisDrone [91] captured in Asia. The conventional approach fails to detect aerial objects from the MAVREC dataset precisely. In contrast, our object detector pretrained on the ground and aerial images from the MAVREC dataset contextualizes the object proposals of that specific geography and enhances the aerial visual perception, thus outperforming other object detectors pre-trained on popular ground-view dataset (MS-COCO [44]) or other aerial datasets collected from different geographies; also, see Figure 5 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0ce972f-e024-4492-877a-e4b1daafd5f2Cited by top-tier papers8
- CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosTrong-Thuan Nguyen, Pha A. Nguyen, Xin Li, Jackson David Cothren et al.NeurIPS 2024 · 13 citations
- Video2BEV: Transforming Drone Videos to BEVs for Video-Based Geo-LocalizationHao Ju, Shaofei Huang, Si Liu, Zhedong ZhengICCV 2025 · 5 citations
- VGGT-Segmentor: Geometry-Enhanced Cross-View SegmentationYulu Gao, Bohao Zhang, Zongheng Tang, Jitong Liao et al.CVPR 2026 · 3 citations
- CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language DetectionZhipeng Liu, Chunbo LuoCVPR 2026 · 1 citation
- LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of LivingDominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind et al.CVPR 2025
Builds on11
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Unbiased Teacher v2: Semi-supervised Object Detection for Anchor-free and Anchor-based DetectorsYen-Cheng Liu, Chih-Yao Ma, Zsolt KiraCVPR 2022 · 124 citations
- Active Teacher for Semi-Supervised Object DetectionPeng Mi, Jianghang Lin, Yiyi Zhou, Yunhang Shen et al.CVPR 2022 · 83 citations
- Dense Learning based Semi-Supervised Object DetectionBinghui Chen, Pengyu Li, Xiang Chen, Biao Wang et al.CVPR 2022 · 80 citations
- MOR-UAV: A Benchmark Dataset and Baselines for Moving Object Recognition in UAV VideosMurari Mandal, Lav Kush Kumar, Santosh Kumar VipparthiACM MM 2020 · 58 citations
Related papers
- MODA: The First Challenging Benchmark for Multispectral Object Detection in Aerial ImagesShuaihao Han, Tingfa Xu, Peifu Liu, Jianan LiAAAI 2026 · 1 citation
- UniGeoRS: A Unified Benchmark for Tri-view Geo-LocalizationXiao Liang, Huaizhi Tang, Feiyang Zhang, Shiji Yuan et al.CVPR 2026
- MMGeo: Multimodal Compositional Geo-Localization for UAVsYuxiang Ji, Boyong He, Zhuoyue Tan, Liaoni WuICCV 2025 · 5 citations
- GeoMIM: Towards Better 3D Knowledge Transfer via Masked Image Modeling for Multi-view 3D UnderstandingJihao Liu, Tai Wang, Boxiao Liu, Qihang Zhang et al.ICCV 2023 · 22 citations
- AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional RelationsJunli Liu, Qizhi Chen, Zhigang Wang, Yiwen Tang et al.ICCV 2025 · 5 citations
