ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical Judgment
Ruochen Li, Jun Li, Bailiang Jian, Kun Yuan, Youxiang Zhu
摘要
Automatically generated radiology reports often receive high scores from existing evaluation metrics but fail to earn clinicians' trust. This gap reveals fundamental flaws in how current metrics assess the quality of generated reports. We rethink the design and evaluation of these metrics and propose a clinically grounded Meta-Evaluation framework. We define clinically grounded criteria spanning clinical alignment and key metric capabilities, including discrimination, robustness, and monotonicity. Using a fine-grained dataset of ground truth and rewritten report pairs annotated with error types, clinical significance labels, and explanations, we systematically evaluate existing metrics and reveal their limitations in interpreting clinical semantics, such as failing to distinguish clinically significant errors, over-penalizing harmless variations, and lacking consistency across error severity levels. Our framework offers guidance for building more clinically reliable evaluation methods. Project link is https: //ruochenli99.github.io/ReEvalMed/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and GenerationSuhyeon Lee, Won Jun Kim, Jinho Chang, Jong Chul YeICLR 2024 · 被引用 80 次
- AlignScore: Evaluating Factual Consistency with A Unified Alignment FunctionYuheng Zha, Yichi Yang, Ruichen Li, Zhiting HuACL 2023 · 被引用 44 次
- RaTEScore: A Metric for Radiology Report GenerationWeike Zhao, Chaoyi Wu, Xiaoman Zhang, Ya Zhang 等EMNLP 2024 · 被引用 17 次
- Alice Benchmarks: Connecting Real World Re-Identification with the SyntheticXiaoxiao Sun, Yue Yao, Shengjin Wang, Hongdong Li 等ICLR 2024 · 被引用 6 次
相关 Paper
- ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report GenerationYunyi Liu, Yingshu Li, Zhanyu Wang, Xinyu Liang 等AAAI 2026
- Image-aware Evaluation of Generated Medical ReportsGefen Dawidowicz, Elad Hirsch, Ayellet TalNeurIPS 2024 · 被引用 3 次
- CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report GenerationRuifeng Yuan, Wanxing Chang, Weiwei Cao, Bowen Shi 等ACL 2026
- GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report EvaluationZhenxuan Zhang, Kinhei Lee, Peiyuan Jing, Weihang Deng 等AAAI 2026 · 被引用 1 次
- CREPE: Rapid Chest X-ray Report Evaluation by Predicting Multi-category Error CountsGihun Cho, Seunghyun Jang, Hanbin Ko, Inhyeok Baek 等EMNLP 2025
