Language-Guided Visual Prompt Compensation for Multi-Modal Remote Sensing Image Classification with Modality Absence
Ling Huang, Wenqian Dong, Song Xiao, Jiahui Qu, Yuanbo Yang, Yunsong Li
摘要
Joint classification of multi-modal remote sensing images has achieved great success thanks to complementary advantages of multi-modal images. However, modality absence is a common dilemma in real world caused by imaging conditions, which leads to a breakdown of most classification methods that rely on complete modalities. Existing approaches either learn shared representations or train specific models for each absence case so that they commonly confront the difficulty of balancing the complementary advantages of the modalities and scalability of the absence case. In this paper, we propose a language-guided visual prompt compensation network (LVPCnet) to achieve joint classification in case of arbitrary modality absence using a unified model that simultaneously considers modality complementarity. It embeds missing modality-specific knowledge into visual prompts to guide the model in capturing complete modal information from available ones for classification. Specifically, a language-guided visual feature decoupling stage (LVFD-stage) is designed to extract shared and specific modal feature from multi-modal images, establishing a complementary representation model of complete modalities. Subsequently, an absence-aware visual prompt compensation stage (VPC-stage) is proposed to learn visual prompts containing missing modality-specific knowledge through cross-modal representation alignment, further guiding the complementary representation model to reconstruct modality-specific features for missing modalities from available ones based on the learned prompts. The proposed VPC-stage entails solely training visual prompts to perceive missing information without retraining the model, facilitating effective scalability to arbitrary modal missing scenarios. Systematic experiments conducted on three public datasets have validated the effectiveness of the proposed approach.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image ParsingXianzhi Ma, Jianhui Li, Changhua Pei, Hao LiuACM MM 2025 · 被引用 3 次
- TAPE: Task-Adaptive Prototype Evolution in Audio-Language Models for Fully Few-shot Class-incremental Audio ClassificationYunlong Gao, Wenxin Liang, Guanglu Wang, Senqi Guan 等CVPR 2026
相关 Paper
- Modal-aware Visual Prompting for Incomplete Multi-modal Brain Tumor SegmentationYansheng Qiu, Ziyuan Zhao, Hongdou Yao, Delin Chen 等ACM MM 2023 · 被引用 25 次
- T-APT: Text-Guided Modality-Aware Prompt Tuning for Arbitrary Multimodal Remote Sensing Data Joint ClassificationQinghao Gao, Jiahui Qu, Wenqian DongAAAI 2026
- SPR: A Structured Prompt Refinement Network for Modality MissingHao Chen, Diwei Su, Zhuo Wang, Zuwang He 等ICML 2026
- Distilled Prompt Learning for Incomplete Multimodal Survival PredictionYingxue Xu, Fengtao Zhou, Chenyu Zhao, Yihui Wang 等CVPR 2025
- Deep Correlated Prompting for Visual Recognition with Missing ModalitiesLianyu Hu, Tongkai Shi, Wei Feng, Fanhua Shang 等NeurIPS 2024 · 被引用 37 次
