Clinical-BERT: Vision-Language Pre-training for Radiograph Diagnosis and Reports Generation
Bin Yan, Mingtao Pei
Abstract
In this paper, we propose a vision-language pre-training model, Clinical-BERT, for the medical domain, and devise three domain-specific tasks: Clinical Diagnosis (CD), Masked MeSH Modeling (MMM), Image-MeSH Matching (IMM), together with one general pre-training task: Masked Language Modeling (MLM), to pre-train the model. The CD task helps the model to learn medical domain knowledge by predicting disease from radiographs. Medical Subject Headings (MeSH) words are important semantic components in radiograph reports, and the MMM task helps the model focus on the prediction of MeSH words. The IMM task helps the model learn the alignment of MeSH words with radiographs by matching scores obtained by a two-level sparse attention: region sparse attention and word sparse attention. Region sparse attention generates corresponding visual features for each word, and word sparse attention enhances the contribution of images-MeSH matching to the matching scores. To the best of our knowledge, this is the first attempt to learn domain knowledge during pre-training for the medical domain. We evaluate the pre-training model on Radiograph Diagnosis and Reports Generation tasks across four challenging datasets: MIMIC-CXR, IU X-Ray, COV-CTR, and NIH, and achieve state-of-the-art results for all the tasks, which demonstrates the effectiveness of our pre-training model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee4378ba-c960-4fee-8548-d7b963274ad5Cited by top-tier papers21
- PromptMRG: Diagnosis-Driven Prompts for Medical Report GenerationHaibo Jin, Haoxuan Che, Yi Lin, Hao ChenAAAI 2024 · 168 citations
- WalkLM: A Uniform Language Model Fine-tuning Framework for Attributed Graph EmbeddingYanchao Tan, Zihao Zhou, Hang Lv, Weiming Liu et al.NeurIPS 2023 · 60 citations
- Automatic Radiology Reports Generation via Memory Alignment NetworkHongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing TianAAAI 2024 · 40 citations
- Continual Self-Supervised Learning: Towards Universal Multi-Modal Medical Data Representation LearningYiwen Ye, Yutong Xie, Jianpeng Zhang, Ziyang Chen et al.CVPR 2024 · 30 citations
- Walking the Tightrope: Autonomous Disentangling Beneficial and Detrimental Drifts in Non-Stationary Custom-TuningXiaoyu Yang, Jie Lu, En YuNeurIPS 2025 · 22 citations
Builds on11
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 552 citations
- VIVO: Visual Vocabulary Pre-Training for Novel Object CaptioningXiaowei Hu, Xi Yin, Kevin Lin, Lei Zhang et al.AAAI 2021 · 63 citations
Related papers
- MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray DiagnosisChaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang et al.ICCV 2023 · 205 citations
- Incorporating medical knowledge in BERT for clinical relation extractionArpita Roy, Shimei PanEMNLP 2021 · 56 citations
- Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed TomographyBowen Shi, Weiwei Cao, Ruifeng Yuan, Wanxing Chang et al.ICML 2026
- A Disease-Aware Dual-Stage Framework for Chest X-ray Report GenerationPuzhen Wu, Hexin Dong, Yi Lin, Yihao Ding et al.AAAI 2026 · 3 citations
- Visual-Textual Attentive Semantic Consistency for Medical Report GenerationYi Zhou, Lei Huang, Tao Zhou, Huazhu Fu et al.ICCV 2021 · 27 citations
