A Prior Instruction Representation Framework for Remote Sensing Image-text Retrieval
Jiancheng Pan, Qing Ma, Cong Bai
摘要
This paper presents a prior instruction representation framework (PIR) for remote sensing image-text retrieval, aimed at remote sensing vision-language understanding tasks to solve the semantic noise problem. Our highlight is the proposal of a paradigm that draws on prior knowledge to instruct adaptive learning of vision and text representations. Concretely, two progressive attention encoder (PAE) structures, Spatial-PAE and Temporal-PAE, are proposed to perform long-range dependency modeling to enhance key feature representation. In vision representation, Vision Instruction Representation (VIR) based on Spatial-PAE exploits the prior-guided knowledge of the remote sensing scene recognition by building a belief matrix to select key features for reducing the impact of semantic noise. In text representation, Language Cycle Attention (LCA) based on Temporal-PAE uses the previous time step to cyclically activate the current time step to enhance text representation capability. A cluster-wise affiliation loss is proposed to constrain the inter-classes and to reduce the semantic confusion zones in the common subspace. Comprehensive experiments demonstrate that using prior knowledge instruction could enhance vision and text representations and could outperform the state-of-the-art methods on two benchmark datasets, RSICD and RSITMD. Codes are available at https://github.com/Zjut-MultimediaPlus/PIR-pytorch.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper9
- UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain AdaptationSiru Zhong, Xixuan Hao, Yibo Yan, Ying Zhang 等ACM MM 2024 · 被引用 8 次
- Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text RetrievalYang Du, Yuqi Liu, Qin JinACM MM 2024 · 被引用 4 次
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part SegmentationYang Miao, Jan-Nico Zaech, Xi Wang, Fabien Despinoy 等NeurIPS 2025 · 被引用 3 次
- Robust Remote Sensing Image–Text Retrieval with Noisy Correspondenceqiya song, Yiqiang Xie, Yuan Sun, Renwei Dian 等CVPR 2026 · 被引用 1 次
- V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object CorrespondenceJiancheng Pan, Runze Wang, Tianwen Qian, Mohammad Mahdi 等CVPR 2026
相关 Paper
- Entity-Level Alignment with Prompt-Guided Adapter for Remote Sensing Image-Text RetrievalShuoshuo Li, Shuli Cheng, Liejun WangACM MM 2025 · 被引用 2 次
- PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text RetrievalPengxiang Ouyang, Qing Ma, Zheng Wang, Cong BaiAAAI 2026
- Accurate and Lightweight Learning for Specific Domain Image-Text RetrievalRui Yang, Shuang Wang, Jianwei Tao, Yingping Han 等ACM MM 2024 · 被引用 6 次
- ZoRI: Towards Discriminative Zero-Shot Remote Sensing Instance SegmentationShiqi Huang, Shuting He, Bihan WenAAAI 2025
- Selection and Reconstruction of Key Locals: A Novel Specific Domain Image-Text Retrieval MethodYu Liao, Xinfeng Zhang, Rui Yang, Jianwei Tao 等ACM MM 2024 · 被引用 2 次
