MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive Learning
Zhe Li, Laurence T. Yang, Bocheng Ren, Xin Nie, Zhangyang Gao, Cheng Tan, Stan Z. Li
Abstract
The scarcity of annotated data has sparked significant interest in unsupervised pre-training methods that leverage medical reports as auxiliary signals for medical visual representation learning. However, existing research overlooks the multi-granularity nature of medical visual representation and lacks suitable contrastive learning techniques to improve the models' generalizability across different granularities, leading to the underutilization of image-text information. To address this, we propose MLIP, a novel framework leveraging domain-specific medical knowledge as guiding signals to integrate language information into the visual domain through imagetext contrastive learning. Our model includes global contrastive learning with our designed divergence encoder, local token-knowledge-patch alignment contrastive learning, and knowledge-guided category-level contrastive learning with expert knowledge. Experimental evaluations reveal the efficacy of our model in enhancing transfer performance for tasks such as image classification, object detection, and semantic segmentation. Notably, MLIP surpasses state-ofthe-art methods even with limited annotated data, highlighting the potential of multimodal pre-training in advancing medical representation learning. 1 * Corresponding Author. 1 Codes are available at https://github.com/gentlefress/MLIP The lung volumes remain low. Signs of mild fluid overload and the extent of the known left pleural effusion have improved.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 784876d8-2a27-4bc1-a653-427f56b9f10bCited by top-tier papers7
- From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent GuidanceZhe Li, Yangyang Wei, Boan Zhu, Yibo Peng et al.ICLR 2026 · 29 citations
- AFiRe: Anatomy-Driven Self-Supervised Learning for Fine-Grained Representation in Radiographic ImagesYihang Liu, Lianghua He, Ying Wen, Longzhen Yang et al.AAAI 2025 · 7 citations
- LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited PairingHuimin Yan, Liang Bai, Xian Yang, Long ChenICML 2026 · 1 citation
- Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report GenerationLongzhen Yang, Zhangkai Ni, Ying Wen, Yihang Liu et al.ACM MM 2025
- KAMP: Knowledge-Anchored Multimodal Pretraining Framework for Medical Image RepresentationFeiyu Huang, Jia Li, Zhao Chen, Yang Wu et al.CVPR 2026
Builds on21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- K-BERT: Enabling Language Representation with Knowledge GraphWeijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang et al.AAAI 2020 · 898 citations
Related papers
- Boosting Medical Visual Understanding From Multi-Granular Language LearningZihan Li, Yiqing Wang, Sina Farsiu, Paul KinahanICLR 2026 · 6 citations
- MMCLIP: Cross-Modal Attention Masked Modelling for Medical Language-Image Pre-TrainingBiao Wu, Yutong Xie, Zeyu Zhang, Vu Minh Hieu Phan et al.ACL 2026 · 4 citations
- Align, Reason and Learn: Enhancing Medical Vision-and-Language Pre-training with KnowledgeZhihong Chen, Guanbin Li, Xiang WanACM MM 2022 · 82 citations
- CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language AlignmentSajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo et al.CVPR 2024
- Towards Medical Vision-Language Contrastive Pre-training via Study-Oriented Semantic ExplorationBo Liu, Zexin Lu, Yan WangACM MM 2024 · 8 citations
