BiomedCoOp: Learning to Prompt for Biomedical Vision-Language Models
Taha Koleilat, Hojat Asgariandehkordi, Hassan Rivaz, Yiming Xiao
Abstract
Recent advancements in vision-language models (VLMs), such as CLIP, have demonstrated substantial success in self-supervised representation learning for vision tasks. However, effectively adapting VLMs to downstream applications remains challenging, as their accuracy often depends on time-intensive and expertise-demanding prompt engineering, while full model fine-tuning is costly. This is particularly true for biomedical images, which, unlike natural images, typically suffer from limited annotated datasets, unintuitive image contrasts, and nuanced visual features. Recent prompt learning techniques, such as Context Optimization (CoOp) intend to tackle these issues, but still fall short in generalizability. Meanwhile, explorations in prompt learning for biomedical image analysis are still highly limited. In this work, we propose BiomedCoOp, a novel prompt learning framework that enables efficient adaptation of BiomedCLIP for accurate and highly generalizable few-shot biomedical image classification. Our approach achieves effective prompt context learning by leveraging semantic consistency with average prompt ensembles from Large Language Models (LLMs) and knowledge distillation with a statistics-based prompt selection strategy. We conducted comprehensive validation of our proposed framework on 11 medical datasets across 9 modalities and 10 organs against existing state-of-the-art methods, demonstrating significant improvements in both accuracy and generalizability. The code is publicly available at https://github.com/HealthX-Lab/BiomedCoOp . * Corresponding author training (CLIP) [37] , which align visual and textual information through contrastive pre-training, allow the exploration of open-set visual concepts, thanks to the adoption of natural language supervision. However, the success of these models often relies heavily on the quality of the textual prompts that guide their predictions while full-model fine-tuning for large-scale VLMs is impractical. To mitigate these, prompt learning that optimizes textual prompts in vision-language models [25, 50, 51] has emerged as one of the critical techniques to enhance performance without the need for extensive fine-tuning. Notably, the pioneering work of Context Optimization (CoOp) [51] introduced this approach for CLIP by treating text prompts as learnable context vectors and preserving the pre-trained model weights. Meanwhile, other approaches [16, 19, 47] focus on lightweight few-shot adaptation through Adapters [18] and Linear Probes [37] to offer parameter-efficient solutions for model adaptation in downstream tasks. Different from natural images, biomedical images include a wide range of contrasts and modalities, depending on the image acquisition devices and parameters. These images, such as MRI and ultrasound, often have unique visual appearances that can be more difficult to interpret than typical photographs. In addition, image features (e.g., color, texture, shape, and anatomical context) that are related to physiological and pathological changes are more nuanced and complex to describe, and can differ between image modalities. Finally, due to privacy concerns and the high requirement for clinical expertise, large datasets of wellannotated biomedical images are scarce for developing clinical deep learning models. While VLMs and the associated prompt learning techniques have shown success across natural image datasets and benchmarks, their application in the biomedical imaging domain (e.g., diagnosis), which has distinct challenges, remains largely under-explored. Due to the unique domain knowledge of biomedical images, the backbone vision-language model for prompt learning may require tailored pre-training for the best outcome. Biomed-specific VLMs, such as BiomedCLIP [48]-pretrained on 15 million biomedical image-text pairs from internet resources-are better suited for biomedical tasks This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fec70ea1-39c2-4c9e-87e6-9858eb4dcdf8Cited by top-tier papers11
- MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image SegmentationTaha Koleilat, Hojat Asgariandehkordi, Omid Nejatimanzari, Berardino Barile et al.CVPR 2026 · 4 citations
- MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQAHaowen Gu, Gensheng Pei, Zeren Sun, Mingwu Ren et al.CVPR 2026 · 2 citations
- SoC: Semantic Orthogonal Calibration for Test-Time Prompt TuningLeo Fillioux, Omprakash Chakraborty, Ismail Ben Ayed, Paul-Henry Cournède et al.CVPR 2026 · 2 citations
- Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language ModelsJiaxiang Liu, Jiawei Du, Xiao Liu, Shangyang Li et al.ICML 2026 · 2 citations
- Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active LearningZhuofan Xie, Zishan Lin, Jinliang Lin, Jie Qi et al.CVPR 2026 · 1 citation
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Prompt-aligned Gradient for Prompt TuningBeier Zhu, Yulei Niu, Yucheng Han, Yue Wu et al.ICCV 2023 · 475 citations
- Self-regulating Prompts: Foundational Model Adaptation without ForgettingMuhammad Uzair Khattak, Syed Talal Wasim, Muzammal Naseer, Salman Khan et al.ICCV 2023 · 365 citations
Related papers
- BioDPP: Dynamic Prompt Policy Learning for Biomedical Vision-Language ModelsPingyi Miao, Xianlai Chen, Kai Sun, Yunbo Wang et al.AAAI 2026
- MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language ModelsYijun Wang, Siying Wu, Lubin Gan, Zheyu Zhang et al.ACM MM 2025
- vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMsMinye Shao, Sihan Guo, Xinrun Li, Xingyu Miao et al.AAAI 2026
- Visual-Language Prompt Tuning with Knowledge-Guided Context OptimizationHantao Yao, Rui Zhang, Changsheng XuCVPR 2023
- MaPLe: Multi-modal Prompt LearningMuhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz, Salman H. Khan et al.CVPR 2023
