ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation
Yunyi Liu, Yingshu Li, Zhanyu Wang, Xinyu Liang, Lingqiao Liu, Lei Wang, Luping Zhou
摘要
Automated radiology report generation (R2Gen) has advanced significantly, yet evaluation remains challenging due to the complexity of assessing report quality. Traditional metrics often misalign with human judgments, failing to identify specific deficiencies. To address this, we introduce ReFINE, a framework for training an Evaluation Model using a novel margin-based reward enforcement loss. This approach decomposes report quality into fine-grained sub-scores across user-defined criteria, improving interpretability. Leveraging GPT-4, we generate diverse training data with paired accepted and rejected reports to train our model under a reward-based system. The trained ReFINE Score provides both granular sub-scores and an aggregated quality assessment, enabling criterion-specific evaluation. Experimental results demonstrate ReFINE's superior alignment with human judgments, outperforming traditional metrics in model selection. Its robustness is validated across three expert-annotated datasets—including chest X-rays and multimodal reports covering 9 imaging modalities—and under two distinct scoring systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 被引用 254 次
- RaTEScore: A Metric for Radiology Report GenerationWeike Zhao, Chaoyi Wu, Xiaoman Zhang, Ya Zhang 等EMNLP 2024 · 被引用 17 次
相关 Paper
- ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical JudgmentRuochen Li, Jun Li, Bailiang Jian, Kun Yuan 等EMNLP 2025
- GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report EvaluationZhenxuan Zhang, Kinhei Lee, Peiyuan Jing, Weihang Deng 等AAAI 2026 · 被引用 1 次
- CREPE: Rapid Chest X-ray Report Evaluation by Predicting Multi-category Error CountsGihun Cho, Seunghyun Jang, Hanbin Ko, Inhyeok Baek 等EMNLP 2025
- Automated Structured Radiology Report GenerationJean-Benoit Delbrouck, Justin Xu, Johannes Moll, Alois Thomas 等ACL 2025
- Exploring the Boundaries of GPT-4 in RadiologyQianchu Liu, Stephanie L. Hyland, Shruthi Bannur, Kenza Bouzid 等EMNLP 2023 · 被引用 22 次
