Entity-Level Alignment with Prompt-Guided Adapter for Remote Sensing Image-Text Retrieval
Shuoshuo Li, Shuli Cheng, Liejun Wang
摘要
Remote Sensing Image-Text Retrieval (RSITR) is a fundamental task in the remote sensing (RS) field and has seen significant progress in recent years. However, existing methods often overlook explicit attention to semantic entities in RS scenes, limiting their capabilities in fine-grained semantic modeling and cross-modal matching, thereby hindering retrieval performance. To address these limitations, we propose a novel framework, Entity-level Alignment with Prompt-guided Adapter (EAPA), which enhances retrieval performance by explicitly perceiving, embedding, and aligning semantic entities in RS images and texts. Built upon the Contrastive Language-Image Pretraining (CLIP) model, EAPA comprises three key modules: the Prompt-guided Attention Adapter (PAA) module, the Pseudo-label-supervised Entity Embedding (PEE) module, and the Cross-modal Entity-level Semantic Alignment (CESA) module. Specifically, PAA freezes the CLIP backbone and introduces learnable prompt vectors to capture RS-specific entity-level semantic knowledge, guiding attention distribution and enhancing semantic representations. To obtain cross-modal consistent entity-level representations, PEE employs an entity query-based encoder to extract entity embeddings of both images and texts, and uses pseudo semantic labels as supervision to ensure that each embedding corresponds to a unique and well-defined semantic category. Based on this, CESA performs one-to-one alignment of cross-modal entity embeddings that correspond to the same semantic category, effectively avoiding mismatches and enhancing fine-grained alignment. Extensive experiments on the RSICD and RSITMD datasets demonstrate that EAPA outperforms state-of-the-art methods across multiple metrics, validating the effectiveness of each module in enhancing fine-grained semantic modeling and cross-modal matching.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CF-IPT: Cross-Modal Fusion Interactive Prompt Tuning of Vision-Language Pre-Trained Model for Multisource Remote Sensing Data ClassificationJinheng Ji, Jiahui Qu, Wenqian Dong, Yunsong LiCVPR 2026
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao 等CVPR 2022 · 被引用 337 次
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su 等AAAI 2025 · 被引用 23 次
- Prompt-Driven Referring Image Segmentation with Instance ContrastingChao Shang, Zichen Song, Heqian Qiu, Lanxiao Wang 等CVPR 2024 · 被引用 20 次
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 被引用 12 次
