AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
Junli Liu, Qizhi Chen, Zhigang Wang, Yiwen Tang, Yiting Zhang, Chi Yan, Dong Wang, Xuelong Li, Bin Zhao
摘要
Visual grounding (VG) aims to localize target objects in an image based on natural language descriptions. In this paper, we propose AerialVG, a new task focusing on visual grounding from aerial views. Compared to traditional VG, AerialVG poses new challenges, e.g., appearance-based grounding is insufficient to distinguish among multiple visually similar objects, and positional relations should be emphasized. Besides, existing VG models struggle when applied to aerial imagery, where high-resolution images cause significant difficulties. To address these challenges, we introduce the first AerialVG dataset, consisting of 5K real-world aerial images, 50K manually annotated descriptions, and 103K objects. Particularly, each annotation in AerialVG dataset contains multiple target objects annotated with relative spatial relations, requiring models to perform comprehensive spatial reasoning. Furthermore, we propose an innovative model especially for the AerialVG task, where a Hierarchical Cross-Attention is devised to focus on target regions, and a Relation-Aware Grounding module is designed to infer positional relations. Experimental results validate the effectiveness of our dataset and method, highlighting the importance of spatial reasoning in aerial visual grounding. The code and dataset will be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- OpenFly: A COMPREHENSIVE PLATFORM FOR AERIAL VISION-LANGUAGE NAVIGATIONYunpeng Gao, Chenhui Li, Zhongrui You, Junli Liu 等ICLR 2026 · 被引用 61 次
- Embodied Crowd CountingRunling Long, Yunlong Wang, Jia Wan, Xiang Deng 等NeurIPS 2025 · 被引用 3 次
- Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And DetectionGuoting Wei, Xia Yuan, Yangzhou, Haizhao Jing 等ICML 2026 · 被引用 2 次
- AerialMind: Towards Referring Multi-Object Tracking in UAV ScenariosChenglizhao Chen, Shaofeng Liang, Runwei Guan, Xiaolou Sun 等AAAI 2026 · 被引用 2 次
它引用的顶会 Paper26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
- Dynamic DETR: End-to-End Object Detection with Dynamic AttentionXiyang Dai, Yinpeng Chen, Jianwei Yang, Pengchuan Zhang 等ICCV 2021 · 被引用 429 次
相关 Paper
- Visual Grounding in Remote Sensing ImagesYuxi Sun, Shanshan Feng, Xutao Li, Yunming Ye 等ACM MM 2022 · 被引用 76 次
- AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based ReferringXinyi Wang, Na Zhao, Zhiyuan Han, Dan Guo 等AAAI 2025 · 被引用 12 次
- ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual GroundingMinghang Zheng, Jiahua Zhang, Qingchao Chen, Yuxin Peng 等ACM MM 2024 · 被引用 5 次
- AerialVLN: Vision-and-Language Navigation for UAVsShubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang 等ICCV 2023 · 被引用 132 次
- Exploiting Contextual Objects and Relations for 3D Visual GroundingLi Yang, Chunfeng Yuan, Ziqi Zhang, Zhongang Qi 等NeurIPS 2023 · 被引用 33 次
