BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
Alejandro Lozano, Min Woo Sun, James Burgess, Liangyu Chen, Jeffrey J. Nirschl, Jeffrey Gu, Ivan Lopez, Josiah Aklilu, Anita Rau, Austin Wolfgang Katzer, Yuhui Zhang, Collin Chiu
Abstract
The development of vision-language models (VLMs) is driven by large-scale and diverse multi-modal datasets. However, progress toward generalist biomedical VLMs is limited by the lack of annotated, publicly accessible datasets across biology and medicine. Existing efforts are limited to narrow domains, missing the full diversity of biomedical knowledge encoded in scientific literature. To address this gap, we introduce BIOMEDICA: a scalable, open-source framework to extract, annotate, and serialize the entirety of the PubMed Central Open Access subset into an easy-to-use, publicly accessible dataset. Our framework produces a comprehensive archive with over 24 million unique image-text pairs from over 6 million articles. Metadata and expert-guided annotations are additionally provided. We demonstrate the utility and accessibility of our resource by releasing BMC-CLIP, a suite of CLIP-style models continuously pre-trained on BIOMEDICA dataset via streaming (eliminating the need to download 27 TB of data locally). On average, our models achieve state-of-theart performance across 40 tasks -spanning pathology, radiology, ophthalmology, dermatology, surgery, molecular biology, parasitology, and cell biology -excelling in zero-shot classification with 6.56% average improvement (as high as 29.8% and 17.5% in dermatology and ophthalmology, respectively) and stronger image-text retrieval while using 10x less compute. To foster reproducibility and collaboration, we release our codebase 1 , 2 and dataset 3 to the broader research community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d22dc4c-6021-49d4-b304-dcdf95cd1375Cited by top-tier papers7
- Can Large Language Models Match the Conclusions of Systematic Reviews?Christopher Polzak, Alejandro Lozano, Min Woo Sun, James Burgess et al.ICLR 2026 · 9 citations
- Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation ModelsZizhi Chen, Yizhen Gao, Minghao Han, Yizhou Liu et al.CVPR 2026 · 3 citations
- Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMsXiaoke Huang, Ningsen Wang, Hui Liu, Xianfeng Tang et al.ICLR 2026 · 3 citations
- CHIPS: Efficient CLIP Adaptation via Curvature-aware Hybrid Influence-based Data SelectionXinlin Zhuang, Yichen Li, Xiwei Liu, Haolin Yang et al.CVPR 2026 · 1 citation
- MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific ResearchJames Burgess, Jeffrey J. Nirschl, Laura Bravo-Sánchez, Alejandro Lozano et al.CVPR 2025
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li et al.CVPR 2022 · 364 citations
- Data Filtering NetworksAlex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt et al.ICLR 2024 · 251 citations
Related papers
- Derm1M: A Million-Scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for DermatologySiyuan Yan, Ming Hu, Yiwen Jiang, Xieji Li et al.ICCV 2025 · 7 citations
- PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent CollaborationYuxuan Sun, Yunlong Zhang, Yixuan Si, Chenglu Zhu et al.ICLR 2025
- MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from TextbooksWenqi Zeng, Yuqi Sun, Chenxi Ma, Weimin Tan et al.ACM MM 2025 · 5 citations
- CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language AlignmentSajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo et al.CVPR 2024
- MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language ModelsYijun Wang, Siying Wu, Lubin Gan, Zheyu Zhang et al.ACM MM 2025
