Divide and Conquer Radiology Report Generation via Observation Level Fine-grained Pretraining and Prompt Tuning
Yuanpin Zhou, Huogen Wang
Abstract
The automation of radiology report generation (RRG) holds immense potential to alleviate radiologists' workloads and improve diagnostic accuracy. Despite advancements in image captioning and vision-language pretraining, RRG remains challenging due to the lengthy and complex nature of radiology reports. In this work, we proposes the Divide and Conquer Radiology Report Generation (DCRRG) model, which breaks down full-text radiology reports into concise observation descriptions. This approach enables the model to capture fine-grained representations from each observation through a two-stage process: an encoding stage focusing on observation prediction tasks to learn fine-grained representations, and a decoding stage for integrating these descriptions into cohesive and comprehensive radiology reports. Experimental results on two benchmark datasets demonstrate that DCRRG achieves significant improvements across all evaluation metrics, underscoring its capability to generate semantically coherent and clinically accurate radiology reports.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2f42750-dbd4-4d78-b6dc-52e3a974db5fCited by top-tier papers1
Ask how each one uses itBuilds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 907 citations
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training ParadigmYangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui et al.ICLR 2022 · 565 citations
Related papers
- Divide and Conquer: Isolating Normal-Abnormal Attributes in Knowledge Graph-Enhanced Radiology Report GenerationXiao Liang, Yanlei Zhang, Di Wang, Haodi Zhong et al.ACM MM 2024 · 7 citations
- Cross-Counter-Repeat Attention for Enhanced Understanding of Visual Semantics in Radiology Report GenerationXiaolei Bo, Feiyang Yang, Feilong Xu, Xiaoli ZhangACM MM 2025
- Bootstrapping Large Language Models for Radiology Report GenerationChang Liu, Yuanhe Tian, Weidong Chen, Yan Song et al.AAAI 2024 · 84 citations
- Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report GenerationKang Liu, Zhuoqi Ma, Xiaolu Kang, Yunan Li et al.CVPR 2025
- A Self-Boosting Framework for Automated Radiographic Report GenerationZhanyu Wang, Luping Zhou, Lei Wang, Xiu LiCVPR 2021
