RIS-LAD: A Benchmark and Model for Referring Image Segmentation in Low-Altitude Drone Imagery
Kai Ye, YingShi Luan, Zhudi Chen, Guangyue Meng, Pingyang Dai, Liujuan Cao
Abstract
Referring Image Segmentation (RIS), which aims to segment specific objects based on natural language descriptions, plays an essential role in vision-language understanding. Despite its progress in remote sensing applications, RIS under Low-Altitude Drone (LAD) scenarios remains underexplored, as existing datasets and methods are typically designed for high-altitude and static-view imagery. They struggled to handle the unique characteristics of LAD views, such as diverse viewpoints and high object density. In this paper, we propose RIS-LAD, the first finegrained RIS benchmark tailored for LAD scenarios, featuring 13,871 meticulously annotated image-text-mask triplets collected from real-world drone footage with emphasis on small, densely cluttered objects and multi-view perspectives. It highlights new challenges not present in previous benchmarks, such as category drift caused by tiny objects and object drift under crowded objects of the same class. Additionally, we propose the Semantic-Aware Adaptive Reasoning Network (SAARN), which decomposes and adaptively routes semantic information to different network stages rather than uniformly injecting all linguistic features. Specifically, the Category-Dominated Linguistic Enhancement (CDLE) aligns visual features with object categories during early encoding, while the Adaptive Reasoning Fusion Module (ARFM) dynamically selects semantic cues across scales to enhance reasoning in complex scenes. Extensive experiments reveal that RIS-LAD presents substantial challenges to state-ofthe-art RIS algorithms, and also demonstrate the effectiveness of our proposed model in addressing these challenges.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51874cce-3963-49a7-ab0c-6ce901dd6792Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.CVPR 2022 · 319 citations
- Visual Grounding in Remote Sensing ImagesYuxi Sun, Shanshan Feng, Xutao Li, Yunming Ye et al.ACM MM 2022 · 76 citations
Related papers
- VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal SegmentationJihwan Hong, Jaeyoung DoCVPR 2026 · 2 citations
- PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning SegmentationShuyan Ke, Yifan Mei, Changli Wu, Yonghan Zheng et al.CVPR 2026 · 3 citations
- Two-stage Visual Cues Enhancement Network for Referring Image SegmentationYang Jiao, Zequn Jie, Weixin Luo, Jingjing Chen et al.ACM MM 2021 · 24 citations
- Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image SegmentationSihan Liu, Yiwei Ma, Xiaoqing Zhang, Haowei Wang et al.CVPR 2024
- Bottom-Up Shift and Reasoning for Referring Image SegmentationSibei Yang, Meng Xia, Guanbin Li, Hong-Yu Zhou et al.CVPR 2021
