Bootstrapping Large Language Models for Radiology Report Generation
Chang Liu, Yuanhe Tian, Weidong Chen, Yan Song, Yongdong Zhang
Abstract
Radiology report generation (RRG) aims to automatically generate a free-text description from a specific clinical radiograph, e.g., chest X-Ray images. Existing approaches tend to perform RRG with specific models trained on the public yet limited data from scratch, where they often lead to inferior performance owing to the problem of inefficient capabilities in both aligning visual and textual features and generating informative reports accordingly. Currently, large language models (LLMs) offered a promising solution to text generation with their power in learning from big data, especially for cross-modal scenarios such as RRG. However, most existing LLMs are pre-trained on general data, and suffer from the same problem of conventional approaches caused by knowledge gap between general and medical domain if they are applied to RRG. Therefore in this paper, we propose an approach to bootstrapping LLMs for RRG with a in-domain instance induction and a coarse-to-fine decoding process. Specifically, the in-domain instance induction process learns to align the LLM to radiology reports from general texts through contrastive learning. The coarse-to-fine decoding performs a text elevating process for those reports from the ranker, further enhanced with visual features and refinement prompts. Experimental results on two prevailing RRG datasets, namely, IU X-Ray and MIMIC-CXR, demonstrate the superiority of our approach to previous state-of-the-art solutions. Further analyses illustrate that, for the LLM, the induction process enables it to better align with the medical domain and the coarse-to-fine generation allows it to conduct more precise text generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e00c0b18-e9c5-4a8e-9143-5a0630f0dd0aCited by top-tier papers21
- Walking the Tightrope: Autonomous Disentangling Beneficial and Detrimental Drifts in Non-Stationary Custom-TuningXiaoyu Yang, Jie Lu, En YuNeurIPS 2025 · 22 citations
- Radiology Report Generation via Multi-objective Preference OptimizationTing Xiao, Lei Shi, Peng Liu, Zhe Wang et al.AAAI 2025 · 21 citations
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-Language Pre-TrainingWeiwei Cao, Jianpeng Zhang, Zhongyi Shui, Sinuo Wang et al.ICCV 2025 · 18 citations
- Dual-path Collaborative Generation Network for Emotional Video CaptioningCheng Ye, Weidong Chen, Jingyu Li, Lei Zhang et al.ACM MM 2024 · 15 citations
- MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and DiagnosisYuting Zhang, Kaishen Yuan, Hao Lu, Yutao Yue et al.CVPR 2026 · 11 citations
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 907 citations
Related papers
- HC-LLM: Historical-Constrained Large Language Models for Radiology Report GenerationTengfei Liu, Jiapu Wang, Yongli Hu, Mingjie Li et al.AAAI 2025 · 6 citations
- Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report GenerationKang Liu, Zhuoqi Ma, Xiaolu Kang, Yunan Li et al.CVPR 2025
- S2D-Align: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report GenerationJiechao Gao, Chang Liu, Yuangang LiAAAI 2026
- Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report GenerationYaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu et al.ICCV 2023 · 47 citations
- LLM-RG4: Flexible and Factual Radiology Report Generation Across Diverse Input ContextsZhuhao Wang, Yihua Sun, Zihan Li, Xuan Yang et al.AAAI 2025 · 6 citations
