A Parameter-Efficient and Fine-Grained Prompt Learning for Vision-Language Models
Yongbin Guo, Shuzhen Li, Zhulin Liu, Tong Zhang, C. L. Philip Chen
Abstract
Current vision-language models (VLMs) understand complex vision-text tasks by extracting overall semantic information from largescale cross-modal associations. However, extracting from large-scale cross-modal associations often smooths out semantic details and requires large computations, limiting multimodal fine-grained understanding performance and efficiency. To address this issue, this paper proposes a detail-oriented prompt learning (DoPL) method for vision-language models to implement fine-grained multi-modal semantic alignment with merely 0.25M trainable parameters. According to the low-entropy information concentration theory, DoPL explores shared interest tokens from text-vision correlations and transforms them into alignment weights to enhance text prompt and vision prompt via detail-oriented prompt generation. It effectively guides the current frozen layer to extract fine-grained text-vision alignment cues. Furthermore, DoPL constructs detail-oriented prompt generation for each frozen layer to implement layer-by-layer localization of finegrained semantic alignment, achieving precise understanding in complex vision-text tasks. DoPL performs well in parameter-efficient finegrained semantic alignment with only 0.12% tunable parameters for vision-language models. The state-of-the-art results over the previous parameter-efficient fine-tuning methods and full fine-tuning approaches on six benchmarks demonstrate the effectiveness and efficiency of DoPL in complex multi-modal tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f56932e4-bfa3-4637-974f-b3c403eaed8fCited by top-tier papers3
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token MergingInha Kang, Youngsun Lim, Seonho Lee, Jiho Choi et al.ICLR 2026 · 1 citation
- Multi-Label Test-Time Adaptation with Bayesian Conditional PriorsQiru Li, Ao Zhou, Zhiwei Jiang, Zifeng Cheng et al.ICML 2026 · 1 citation
- CAPT: Confusion-Aware Prompt Tuning for Reducing Vision-Language MisalignmentMaoyuan Shao, Yutong Gao, Xinyang Huang, Lijuan Sun et al.CVPR 2026 · 1 citation
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
Related papers
- Surrogate Prompt Learning: Towards Efficient and Diverse Prompt Learning for Vision-Language ModelsLiangchen Liu, Nannan Wang, Xi Yang, Xinbo Gao et al.ICML 2025
- CoPL: Contextual Prompt Learning for Vision-Language UnderstandingKoustava Goswami, Srikrishna Karanam, Prateksha Udhayanan, K. J. Joseph et al.AAAI 2024 · 20 citations
- Tuning Multi-mode Token-level Prompt Alignment across ModalitiesDongsheng Wang, Miaoge Li, Xinyang Liu, Mingsheng Xu et al.NeurIPS 2023 · 49 citations
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang et al.NeurIPS 2025 · 7 citations
- MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task LearningYi Xin, Junlong Du, Qiang Wang, Ke Yan et al.AAAI 2024 · 102 citations
