Entity-Level Alignment with Prompt-Guided Adapter for Remote Sensing Image-Text Retrieval
Shuoshuo Li, Shuli Cheng, Liejun Wang
Abstract
Remote Sensing Image-Text Retrieval (RSITR) is a fundamental task in the remote sensing (RS) field and has seen significant progress in recent years. However, existing methods often overlook explicit attention to semantic entities in RS scenes, limiting their capabilities in fine-grained semantic modeling and cross-modal matching, thereby hindering retrieval performance. To address these limitations, we propose a novel framework, Entity-level Alignment with Prompt-guided Adapter (EAPA), which enhances retrieval performance by explicitly perceiving, embedding, and aligning semantic entities in RS images and texts. Built upon the Contrastive Language-Image Pretraining (CLIP) model, EAPA comprises three key modules: the Prompt-guided Attention Adapter (PAA) module, the Pseudo-label-supervised Entity Embedding (PEE) module, and the Cross-modal Entity-level Semantic Alignment (CESA) module. Specifically, PAA freezes the CLIP backbone and introduces learnable prompt vectors to capture RS-specific entity-level semantic knowledge, guiding attention distribution and enhancing semantic representations. To obtain cross-modal consistent entity-level representations, PEE employs an entity query-based encoder to extract entity embeddings of both images and texts, and uses pseudo semantic labels as supervision to ensure that each embedding corresponds to a unique and well-defined semantic category. Based on this, CESA performs one-to-one alignment of cross-modal entity embeddings that correspond to the same semantic category, effectively avoiding mismatches and enhancing fine-grained alignment. Extensive experiments on the RSICD and RSITMD datasets demonstrate that EAPA outperforms state-of-the-art methods across multiple metrics, validating the effectiveness of each module in enhancing fine-grained semantic modeling and cross-modal matching.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c2f34025-c6c5-4635-8bd6-d95fefaaee31Related papers
- CF-IPT: Cross-Modal Fusion Interactive Prompt Tuning of Vision-Language Pre-Trained Model for Multisource Remote Sensing Data ClassificationJinheng Ji, Jiahui Qu, Wenqian Dong, Yunsong LiCVPR 2026
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao et al.CVPR 2022 · 337 citations
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su et al.AAAI 2025 · 23 citations
- Prompt-Driven Referring Image Segmentation with Instance ContrastingChao Shang, Zichen Song, Heqian Qiu, Lanxiao Wang et al.CVPR 2024 · 20 citations
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 12 citations
