CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation
Ruifeng Yuan, Wanxing Chang, Weiwei Cao, Bowen Shi, Zhongyu Wei, Ling Zhang, Jianpeng Zhang
Abstract
The evaluation of generated reports remains a critical challenge in Computed Tomography (CT) report generation, due to the large volume of text, the diversity and complexity of findings, and the presence of fine-grained, disease-oriented attributes. Conventional evaluation metrics offer only coarse measures of lexical overlap or entity matching and fail to reflect the granular diagnostic accuracy required for clinical use. To address this gap, we propose CT-FineBench, a benchmark built from CT-RATE and Merlin to evaluate the fine-grained factual consistency of CT reports, constructed from CT-RATE and Merlin. Our benchmark is constructed through a meticulous, Question-Answering (QA) based process: first, we identify and structure key, finding-specific clinical attributes (e.g., location, size, margin). Second, we systematically transform these attributes into a QA dataset, where questions probe for specific clinical details grounded in gold-standard reports. The evaluation protocol for CT-FineBench involves using this QA dataset to query a machine-generated report and scoring the correctness of the answers. This allows for a comprehensive, interpretable, and clinically-relevant assessment, moving beyond superficial lexical overlap to pinpoint specific clinical errors. Experiments show that CT-FineBench correlates better with expert clinical assessment and is substantially more sensitive to fine-grained factual errors than prior metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ba5a770-8bd1-42c0-b232-34352cd9306fBuilds on2
Related papers
- Automated Structured Radiology Report GenerationJean-Benoit Delbrouck, Justin Xu, Johannes Moll, Alois Thomas et al.ACL 2025
- CREPE: Rapid Chest X-ray Report Evaluation by Predicting Multi-category Error CountsGihun Cho, Seunghyun Jang, Hanbin Ko, Inhyeok Baek et al.EMNLP 2025
- GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report EvaluationZhenxuan Zhang, Kinhei Lee, Peiyuan Jing, Weihang Deng et al.AAAI 2026 · 1 citation
- MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMsZhan Qu, Michael FärberACL 2026 · 1 citation
- Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image UnderstandingZhongyi Shui, Jianpeng Zhang, Weiwei Cao, Sinuo Wang et al.ICLR 2025
