LIMITR: Leveraging Local Information for Medical Image-Text Representation
Gefen Dawidowicz, Elad Hirsch, Ayellet Tal
Abstract
Medical imaging analysis plays a critical role in the diagnosis and treatment of various medical conditions. This paper focuses on chest X-ray images and their corresponding radiological reports. It presents a new model that learns a joint X-ray image & report representation. The model is based on a novel alignment scheme between the visual data and the text, which takes into account both local and global information. Furthermore, the model integrates domain-specific information of two types—lateral images and the consistent visual structure of chest images. Our representation is shown to benefit three types of retrieval tasks: text-image retrieval, class-based retrieval, and phrase-grounding. Our code is publicly available1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Improving fine-grained understanding in image-text pre-trainingIoana Bica, Anastasija Ilic, Matthias Bauer, Goker Erdogan et al.ICML 2024 · 53 citations
- Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report GenerationWenting Chen, Linlin Shen, Jingyang Lin, Jiebo Luo et al.ACL 2024 · 16 citations
- Image-aware Evaluation of Generated Medical ReportsGefen Dawidowicz, Elad Hirsch, Ayellet TalNeurIPS 2024 · 3 citations
- Self-guided Semantic Inspection for Zero-Shot Composed Image RetrievalJingjing Zhang, Lei Zhang, Zheren Fu, Bo Hu et al.CVPR 2026
- FAMDR: Feature-Aligned Multimodal Denoising for Reliable Diagnostic Reconciliation in Medical ImagingXun Liang, Zhiying Li, Hongxun JiangAAAI 2026
Builds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation LearningFuying Wang, Yuyin Zhou, Shujun Wang, Varut Vardhanabhuti et al.NeurIPS 2022 · 302 citations
Related papers
- Automatic Radiology Reports Generation via Memory Alignment NetworkHongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing TianAAAI 2024 · 40 citations
- Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report GenerationYaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu et al.ICCV 2023 · 47 citations
- Medical Report Generation via Multimodal Spatio-Temporal FusionXin Mei, Rui Mao, Xiaoyan Cai, Libin Yang et al.ACM MM 2024 · 8 citations
- Report-Concept Textual-Prompt Learning for Enhancing X-ray DiagnosisXiongjun Zhao, Zhengyu Liu, Fen Liu, Guanting Li et al.ACM MM 2024 · 3 citations
- DART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report GenerationSang-Jun Park, Keun-Soo Heo, Dong-Hee Shin, Young-Han Son et al.CVPR 2025
