ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation
Yunyi Liu, Yingshu Li, Zhanyu Wang, Xinyu Liang, Lingqiao Liu, Lei Wang, Luping Zhou
Abstract
Automated radiology report generation (R2Gen) has advanced significantly, yet evaluation remains challenging due to the complexity of assessing report quality. Traditional metrics often misalign with human judgments, failing to identify specific deficiencies. To address this, we introduce ReFINE, a framework for training an Evaluation Model using a novel margin-based reward enforcement loss. This approach decomposes report quality into fine-grained sub-scores across user-defined criteria, improving interpretability. Leveraging GPT-4, we generate diverse training data with paired accepted and rejected reports to train our model under a reward-based system. The trained ReFINE Score provides both granular sub-scores and an aggregated quality assessment, enabling criterion-specific evaluation. Experimental results demonstrate ReFINE's superior alignment with human judgments, outperforming traditional metrics in model selection. Its robustness is validated across three expert-annotated datasets—including chest X-rays and multimodal reports covering 9 imaging modalities—and under two distinct scoring systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
- RaTEScore: A Metric for Radiology Report GenerationWeike Zhao, Chaoyi Wu, Xiaoman Zhang, Ya Zhang et al.EMNLP 2024 · 17 citations
Related papers
- ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical JudgmentRuochen Li, Jun Li, Bailiang Jian, Kun Yuan et al.EMNLP 2025
- GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report EvaluationZhenxuan Zhang, Kinhei Lee, Peiyuan Jing, Weihang Deng et al.AAAI 2026 · 1 citation
- CREPE: Rapid Chest X-ray Report Evaluation by Predicting Multi-category Error CountsGihun Cho, Seunghyun Jang, Hanbin Ko, Inhyeok Baek et al.EMNLP 2025
- Automated Structured Radiology Report GenerationJean-Benoit Delbrouck, Justin Xu, Johannes Moll, Alois Thomas et al.ACL 2025
- Exploring the Boundaries of GPT-4 in RadiologyQianchu Liu, Stephanie L. Hyland, Shruthi Bannur, Kenza Bouzid et al.EMNLP 2023 · 22 citations
