GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-efficient Medical Image Recognition
Shih-Cheng Huang, Liyue Shen, Matthew P. Lungren, Serena Yeung
Abstract
In recent years, the growing utilization of medical imaging is placing an increasing burden on radiologists. Deep learning provides a promising solution for automatic medical image analysis and clinical decision support. However, large-scale manually labeled datasets required for training deep neural networks are difficult and expensive to obtain for medical images. The purpose of this work is to develop label-efficient multimodal medical imaging representations by leveraging radiology reports. We propose an attention-based framework for learning global and local representations by contrasting image sub-regions and words in the paired report. In addition, we propose methods to leverage the learned representations for various downstream medical image recognition tasks with limited labels. Our results demonstrate high-performance and label-efficiency for image-text retrieval, classification (finetuning and zerosshot settings), and segmentation on different datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27bd4f79-213d-4699-a402-4cd34c12ddddCited by top-tier papers63
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 907 citations
- Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation LearningFuying Wang, Yuyin Zhou, Shujun Wang, Varut Vardhanabhuti et al.NeurIPS 2022 · 302 citations
- MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray DiagnosisChaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang et al.ICCV 2023 · 205 citations
- Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing BiasZhongwei Wan, Che Liu, Mi Zhang, Jie Fu et al.NeurIPS 2023 · 114 citations
- PRIOR: Prototype Representation Joint Learning from Medical Images and ReportsPujin Cheng, Li Lin, Junyan Lyu, Yijin Huang et al.ICCV 2023 · 91 citations
Builds on5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- VirTex: Learning Visual Representations From Textual AnnotationsKaran Desai, Justin JohnsonCVPR 2021
- IMRAM: Iterative Matching With Recurrent Attention Memory for Cross-Modal Image-Text RetrievalHui Chen, Guiguang Ding, Xudong Liu, Zijia Lin et al.CVPR 2020
Related papers
- Towards Medical Vision-Language Contrastive Pre-training via Study-Oriented Semantic ExplorationBo Liu, Zexin Lu, Yan WangACM MM 2024 · 8 citations
- LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited PairingHuimin Yan, Liang Bai, Xian Yang, Long ChenICML 2026 · 1 citation
- LIMITR: Leveraging Local Information for Medical Image-Text RepresentationGefen Dawidowicz, Elad Hirsch, Ayellet TalICCV 2023 · 27 citations
- CARZero: Cross-Attention Alignment for Radiology Zero-Shot ClassificationHaoran Lai, Qingsong Yao, Zihang Jiang, Rongsheng Wang et al.CVPR 2024
- Report-Concept Textual-Prompt Learning for Enhancing X-ray DiagnosisXiongjun Zhao, Zhengyu Liu, Fen Liu, Guanting Li et al.ACM MM 2024 · 3 citations
