VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object Detection
Shuohao Shi, Qiang Fang, Xin Xu
Abstract
Closed-set object detection in remote sensing imagery has made significant progress, but achieving high detection accuracy remains challenging. Vision-Language Models (VLMs), which possess rich prior knowledge, offer a promising solution to this challenge. However, most existing VLMs are designed for open-vocabulary tasks and exhibit inherent limitations when directly applied to closedset scenarios, such as notable accuracy degradation and high deployment costs. To address these issues, we propose VLM4RSDet, a novel collaborative training framework that leverages vision-language model to enhance the performance of conventional closed-set remote sensing object detectors. Notably, during inference, VLM4RSDet only retains the standard object detection architecture, thus avoiding any additional deployment overhead. Furthermore, we introduce a Global-Local Cross-Attention (GLCA) module and a Learnable Hierarchical Prediction Strategy (LHPS) to further improve collaborative training performance. Extensive experiments on five benchmark datasets demonstrate the effectiveness and robustness of our approach. In particular, our method outperforms the state-of-the-art by 7.5% in mAP 0.5:0.95 on the VisDrone2019 dataset. Our code is available at https://github.com/cszzshi/VLM4RSDet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5950fc18-4e44-4475-9c44-beb4928ae52dBuilds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Oriented R-CNN for Object DetectionXingxing Xie, Gong Cheng, Jiabao Wang, Xiwen Yao et al.ICCV 2021 · 1,070 citations
Related papers
- Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object DetectionYoujun Zhao, Jiaying Lin, Rynson W. H. LauAAAI 2025 · 2 citations
- AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object DetectionHaolin Li, Yaohua Wang, Ze Yan, Lijie Wen et al.CVPR 2026
- LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained DescriptorsSheng Jin, Xueying Jiang, Jiaxing Huang, Lewei Lu et al.ICLR 2024 · 48 citations
- Simple Image-Level Classification Improves Open-Vocabulary Object DetectionRuohuan Fang, Guansong Pang, Xiao BaiAAAI 2024 · 26 citations
- Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted PretrainingYuxuan Li, Yuming Chen, Yunheng Li, Ming-Ming Cheng et al.ICML 2026
