MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray Diagnosis
Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, Weidi Xie
Abstract
In this paper, we consider enhancing medical visual-language pre-training (VLP) with domain-specific knowledge, by exploiting the paired image-text reports from the radiological daily practice. In particular, we make the following contributions: First, unlike existing works that directly process the raw reports, we adopt a novel triplet extraction module to extract the medical-related information, avoiding unnecessary complexity from language grammar and enhancing the supervision signals; Second, we propose a novel triplet encoding module with entity translation by querying a knowledge base, to exploit the rich domain knowledge in medical field, and implicitly build relationships between medical entities in the language embedding space; Third, we propose to use a Transformer-based fusion model for spatially aligning the entity description with visual signals at the image patch level, enabling the ability for medical diagnosis; Fourth, we conduct thorough experiments to validate the effectiveness of our architecture, and benchmark on numerous public benchmarks e.g., ChestX-ray14, RSNA Pneumonia, SIIM-ACR Pneumothorax, COVIDx CXR-2, COVID Rural, and EdemaSeverity. In both zero-shot and fine-tuning settings, our model has demonstrated strong performance compared with the former methods on disease classification and grounding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5fff96c-0ff3-4432-94c9-403ee9ce2a4eCited by top-tier papers23
- Walking the Tightrope: Autonomous Disentangling Beneficial and Detrimental Drifts in Non-Stationary Custom-TuningXiaoyu Yang, Jie Lu, En YuNeurIPS 2025 · 22 citations
- Knowledge-Empowered Dynamic Graph Network for Irregularly Sampled Medical Time SeriesYicheng Luo, Zhen Liu, Linghao Wang, Binquan Wu et al.NeurIPS 2024 · 19 citations
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-Language Pre-TrainingWeiwei Cao, Jianpeng Zhang, Zhongyi Shui, Sinuo Wang et al.ICCV 2025 · 18 citations
- ProtCLIP: Function-Informed Protein Multi-Modal LearningHanjing Zhou, Mingze Yin, Wei Wu, Mingyang Li et al.AAAI 2025 · 11 citations
- From Pretraining to Pathology: How Noise Leads to Catastrophic Inheritance in Medical ModelsHao Sun, Zhongyi Han, Hao Chen, Jindong Wang et al.NeurIPS 2025 · 4 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed TomographyBowen Shi, Weiwei Cao, Ruifeng Yuan, Wanxing Chang et al.ICML 2026
- Align, Reason and Learn: Enhancing Medical Vision-and-Language Pre-training with KnowledgeZhihong Chen, Guanbin Li, Xiang WanACM MM 2022 · 82 citations
- Clinical-BERT: Vision-Language Pre-training for Radiograph Diagnosis and Reports GenerationBin Yan, Mingtao PeiAAAI 2022 · 138 citations
- Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-Training FrameworkVu Minh Hieu Phan, Yutong Xie, Yuankai Qi, Lingqiao Liu et al.CVPR 2024 · 17 citations
- Report-Concept Textual-Prompt Learning for Enhancing X-ray DiagnosisXiongjun Zhao, Zhengyu Liu, Fen Liu, Guanting Li et al.ACM MM 2024 · 3 citations
