Hierarchical Cross-Modal Prompt Learning for Vision-Language Models
Hao Zheng, Shunzhi Yang, Zhuoxin He, Jinfeng Yang, Zhenhua Huang
Abstract
Pre-trained Vision-Language Models (VLMs) such as CLIP have shown excellent generalization abilities. However, adapting these large-scale models to downstream tasks while preserving their generalization capabilities remains challenging. Although prompt learning methods have shown promise, they suffer from two fundamental bottlenecks that limit generalization: (a) modality isolation, and (b) hierarchical semantic decay. To address these limitations, we propose HiCroPL, a Hierarchical Cross-modal Prompt Learning framework that establishes bidirectional knowledge flow between text and vision modalities, enabling them to refine their semantics mutually. HiCroPL routes knowledge flows by leveraging the complementary strengths of text and vision. In early layers, text prompts inject relatively clear semantics into visual prompts through a hierarchical knowledge mapper, enhancing the representation of low-level visual semantics. In later layers, visual prompts encoding specific task-relevant objects flow back to refine text prompts, enabling deeper alignment. Crucially, our hierarchical knowledge mapper allows representations at multi-scales to be fused, ensuring that deeper representations retain transferable shallow semantics thereby enhancing generalization. We further introduce a lightweight layer-specific knowledge proxy to enable efficient cross-modal interactions. Extensive evaluations across four tasks demonstrate HiCroPL's superior performance, achieving state-of-the-art results on 11 benchmarks with significant improvements. Code is available at: https://github.com/zzeoZheng/HiCroPL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext debbf379-03e8-4ee2-bcb7-359e6b1baa1fCited by top-tier papers9
- Exploring Interpretability for Visual Prompt Tuning with Cross-layer ConceptsYubin Wang, Xinyang Jiang, De Cheng, Xiangqian Zhao et al.ICLR 2026 · 1 citation
- Heterogeneous Complementary DistillationLiuchi Xu, Hao Zheng, Lu Wang, Lisheng Xu et al.AAAI 2026
- SpecPL: Disentangling Spectral Granularity for Prompt LearningJingtao Zhou, Xirui Kang, feiyang huang, Lai Man PoICML 2026
- Leveraging Evidence Priors for Robust Prompt Learning under Noisy Supervision in Vision-Language ModelsJunnan Zou, Zhu Teng, Wei Zhang, Ming He et al.ICML 2026
- Prompt-Robust Vision-Language Models via Meta-FinetuningHaohui Liang, Runlin Huang, Yingjun Du, Yujia Hu et al.ICLR 2026
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
Related papers
- MaPLe: Multi-modal Prompt LearningMuhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz, Salman H. Khan et al.CVPR 2023
- COMMA: Co-articulated Multi-Modal LearningLianyu Hu, Liqing Gao, Zekang Liu, Chi-Man Pun et al.AAAI 2024 · 7 citations
- APoLLo : Unified Adapter and Prompt Learning for Vision Language ModelsSanjoy Chowdhury, Sayan Nag, Dinesh ManochaEMNLP 2023 · 17 citations
- Concept-Guided Prompt Learning for Generalization in Vision-Language ModelsYi Zhang, Ce Zhang, Ke Yu, Yushun Tang et al.AAAI 2024 · 37 citations
- HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware AttentionShijie Geng, Jianbo Yuan, Yu Tian, Yuxiao Chen et al.ICLR 2023 · 11 citations
