BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
Alejandro Lozano, Min Woo Sun, James Burgess, Liangyu Chen, Jeffrey J. Nirschl, Jeffrey Gu, Ivan Lopez, Josiah Aklilu, Anita Rau, Austin Wolfgang Katzer, Yuhui Zhang, Collin Chiu
摘要
The development of vision-language models (VLMs) is driven by large-scale and diverse multi-modal datasets. However, progress toward generalist biomedical VLMs is limited by the lack of annotated, publicly accessible datasets across biology and medicine. Existing efforts are limited to narrow domains, missing the full diversity of biomedical knowledge encoded in scientific literature. To address this gap, we introduce BIOMEDICA: a scalable, open-source framework to extract, annotate, and serialize the entirety of the PubMed Central Open Access subset into an easy-to-use, publicly accessible dataset. Our framework produces a comprehensive archive with over 24 million unique image-text pairs from over 6 million articles. Metadata and expert-guided annotations are additionally provided. We demonstrate the utility and accessibility of our resource by releasing BMC-CLIP, a suite of CLIP-style models continuously pre-trained on BIOMEDICA dataset via streaming (eliminating the need to download 27 TB of data locally). On average, our models achieve state-of-theart performance across 40 tasks -spanning pathology, radiology, ophthalmology, dermatology, surgery, molecular biology, parasitology, and cell biology -excelling in zero-shot classification with 6.56% average improvement (as high as 29.8% and 17.5% in dermatology and ophthalmology, respectively) and stronger image-text retrieval while using 10x less compute. To foster reproducibility and collaboration, we release our codebase 1 , 2 and dataset 3 to the broader research community.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Can Large Language Models Match the Conclusions of Systematic Reviews?Christopher Polzak, Alejandro Lozano, Min Woo Sun, James Burgess 等ICLR 2026 · 被引用 9 次
- Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation ModelsZizhi Chen, Yizhen Gao, Minghao Han, Yizhou Liu 等CVPR 2026 · 被引用 3 次
- Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMsXiaoke Huang, Ningsen Wang, Hui Liu, Xianfeng Tang 等ICLR 2026 · 被引用 3 次
- CHIPS: Efficient CLIP Adaptation via Curvature-aware Hybrid Influence-based Data SelectionXinlin Zhuang, Yichen Li, Xiwei Liu, Haolin Yang 等CVPR 2026 · 被引用 1 次
- MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific ResearchJames Burgess, Jeffrey J. Nirschl, Laura Bravo-Sánchez, Alejandro Lozano 等CVPR 2025
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li 等CVPR 2022 · 被引用 364 次
- Data Filtering NetworksAlex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt 等ICLR 2024 · 被引用 251 次
相关 Paper
- Derm1M: A Million-Scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for DermatologySiyuan Yan, Ming Hu, Yiwen Jiang, Xieji Li 等ICCV 2025 · 被引用 7 次
- PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent CollaborationYuxuan Sun, Yunlong Zhang, Yixuan Si, Chenglu Zhu 等ICLR 2025
- MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from TextbooksWenqi Zeng, Yuqi Sun, Chenxi Ma, Weimin Tan 等ACM MM 2025 · 被引用 5 次
- CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language AlignmentSajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo 等CVPR 2024
- MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language ModelsYijun Wang, Siying Wu, Lubin Gan, Zheyu Zhang 等ACM MM 2025
