LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation
Suhyeon Lee, Won Jun Kim, Jinho Chang, Jong Chul Ye
Abstract
Following the impressive development of LLMs, vision-language alignment in LLMs is actively being researched to enable multimodal reasoning and visual input/output. This direction of research is particularly relevant to medical imaging because accurate medical image analysis and generation consist of reasoning based on a combination of visual features and prior knowledge. Many recent works have focused on training adapter networks that serve as an information bridge between image processing (encoding or generating) networks and LLMs; but presumably, in order to achieve maximum reasoning potential of LLMs on visual information as well as language, image and text features should be allowed to interact more freely. This is especially important in the medical domain because understanding and generating medical images such as chest X-rays (CXR) require not only accurate visual and language-based reasoning but also a more intimate mapping between the two modalities. Thus, taking inspiration from previous work on the transformer and VQ-GAN combination for bidirectional image and text generation, we build upon this approach and develop a method for instruction-tuning an LLM pre-trained only on text to gain vision-language capabilities for medical images. Specifically, we leverage a pretrained LLM's existing questionanswering and instruction-following abilities to teach it to understand visual inputs by instructing it to answer questions about image inputs and, symmetrically, output both text and image responses appropriate to a given query by tuning the LLM with diverse tasks that encompass image-based text-generation and text-based image-generation. We show that our model, LLM-CXR, trained in this approach shows better image-text alignment in both CXR understanding and generation tasks while being smaller in size compared to previously developed models that perform a narrower range of tasks. * These authors contributed equally to this work. † In order to comply with the MIMIC-CXR data usage license (Johnson et al., 2019a), all CXR images presented in the Figures 2, 3 , 4, 8 are replaced with similar CXR's from the Indiana University chest X-ray dataset (Demner-Fushman et al., 2016) ; and the presented MIMIC text reports are paraphrased.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a5cb11f-1f7b-4e99-89a6-5a5d2ebb9944Cited by top-tier papers14
- MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and DiagnosisYuting Zhang, Kaishen Yuan, Hao Lu, Yutao Yue et al.CVPR 2026 · 11 citations
- GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-Ray DiagnosisBo Liu, Ke Zou, Li-Ming Zhan, Zexin Lu et al.ICCV 2025 · 10 citations
- InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction DetectionJinguo Luo, Weihong Ren, Quanlong Zheng, Yanhao Zhang et al.NeurIPS 2025 · 3 citations
- Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation DataXun Zhu, Fanbin Mo, Zheng Zhang, Jiaxi Wang et al.ACM MM 2025
- ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical JudgmentRuochen Li, Jun Li, Bailiang Jian, Kun Yuan et al.EMNLP 2025
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem SolvingZixian Guo, Ming Liu, Qilong Wang, Zhilong Ji et al.ICCV 2025 · 1 citation
- Versatile Vision-Language Model for 3D Computed TomographyJiayu Lei, Ziqing Fan, Yanyong Zhang, Weidi Xie et al.AAAI 2026
- MetaMorph: Multimodal Understanding and Generation via Instruction TuningShengbang Tong, David Fan, Jiachen Zhu, Yunyang Xiong et al.ICCV 2025 · 14 citations
- AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray InterpretationQingqiu Li, Zihang Cui, Seongsu Bae, Jilan Xu et al.NeurIPS 2025 · 10 citations
- Medical Vision-Language Pretraining with LLM-Guided Temporal SupervisionLiang Bai, Zhi Wang, Huimin Yan, Xian YangAAAI 2026
