PRIOR: Prototype Representation Joint Learning from Medical Images and Reports
Pujin Cheng, Li Lin, Junyan Lyu, Yijin Huang, Wenhan Luo, Xiaoying Tang
Abstract
Contrastive learning based vision-language joint pre-training has emerged as a successful representation learning strategy. In this paper, we present a prototype representation learning framework incorporating both global and local alignment between medical images and reports. In contrast to standard global multi-modality alignment methods, we employ a local alignment module for fine-grained representation. Furthermore, a cross-modality conditional reconstruction module is designed to interchange information across modalities in the training phase by reconstructing masked images and reports. For reconstructing long reports, a sentence-wise prototype memory bank is constructed, enabling the network to focus on low-level localized visual and high-level clinical linguistic features. Additionally, a non-auto-regressive generation paradigm is proposed for reconstructing non-sequential reports. Experimental results on five downstream tasks, including supervised classification, zero-shot classification, image-to-text retrieval, semantic segmentation, and object detection, show the proposed method outperforms other state-of-the-art methods across multiple datasets and under different dataset size settings. The code is available at https://github.com/QtacierP/PRIOR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd97e966-02fa-4384-9a62-2374d1fe606dCited by top-tier papers18
- Eye-gaze Guided Multi-modal Alignment for Medical Representation LearningChong Ma, Hanqi Jiang, Wenting Chen, Yiwei Li et al.NeurIPS 2024 · 29 citations
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang et al.EMNLP 2024 · 28 citations
- G2D: From Global to Dense Radiography Representation Learning via Vision-Language Pre-trainingChe Liu, Cheng Ouyang, Sibo Cheng, Anand Shah et al.NeurIPS 2024 · 21 citations
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-Language Pre-TrainingWeiwei Cao, Jianpeng Zhang, Zhongyi Shui, Sinuo Wang et al.ICCV 2025 · 18 citations
- Unlocking the Power of Spatial and Temporal Information in Medical Multimodal Pre-trainingJinxia Yang, Bing Su, Xin Zhao, Ji-Rong WenICML 2024 · 13 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Towards Medical Vision-Language Contrastive Pre-training via Study-Oriented Semantic ExplorationBo Liu, Zexin Lu, Yan WangACM MM 2024 · 8 citations
- Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed TomographyBowen Shi, Weiwei Cao, Ruifeng Yuan, Wanxing Chang et al.ICML 2026
- Pathology-Aware Reconstruction with Discriminative Knowledge Boosting Alignment for Che-Xray Vision-Language Pre-trainingLihong Qiao, Shiyi Gao, Yucheng Shu, Bin Xiao et al.ACM MM 2025
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su et al.AAAI 2025 · 23 citations
- LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited PairingHuimin Yan, Liang Bai, Xian Yang, Long ChenICML 2026 · 1 citation
