DocLens: Multi-aspect Fine-grained Medical Text Evaluation
Yiqing Xie, Sheng Zhang, Hao Cheng, Pengfei Liu, Zelalem Gero, Cliff Wong, Tristan Naumann, Hoifung Poon, Carolyn P. Rosé
Abstract
Medical text generation aims to assist with administrative work and highlight salient information to support decision-making. To reflect the specific requirements of medical text, in this paper, we propose a set of metrics to evaluate the completeness, conciseness, and attribution of the generated text at a fine-grained level. The metrics can be computed by various types of evaluators including instruction-following (both proprietary and open-source) and supervised entailment models. We demonstrate the effectiveness of the resulting framework, DO-CLENS, with three evaluators on three tasks: clinical note generation, radiology report summarization, and patient question summarization. A comprehensive human study shows that DOCLENS exhibits substantially higher agreement with the judgments of medical experts than existing metrics. The results also highlight the need to improve open-source evaluators and suggest potential directions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic EvaluationXiangxu Zhang, Lei Li, Yanyun Zhou, Xiao Zhou et al.ACL 2026 · 3 citations
- When KV Cache Reuse Fails in Multi-Agent Systems: Cross-Candidate Interaction is Crucial for LLM JudgesSichu Liang, Zhenglin Wang, Jiajia Chu, Pengfei Xia et al.ACL 2026
- TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical DomainBohao Chu, Meijie Li, Sameh Frihat, Chengyu Gu et al.EMNLP 2025
Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology ReportsYuhao Zhang, Derek Merck, Emily Bao Tsai, Christopher D. Manning et al.ACL 2020 · 160 citations
Related papers
- CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report GenerationRuifeng Yuan, Wanxing Chang, Weiwei Cao, Bowen Shi et al.ACL 2026
- Human Evaluation and Correlation with Automatic Metrics in Consultation Note GenerationFrancesco Moramarco, Alex Papadopoulos-Korfiatis, Mark Perera, Damir Juric et al.ACL 2022 · 62 citations
- NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference ChecklistIftitahu Ni'mah, Meng Fang, Vlado Menkovski, Mykola PechenizkiyACL 2023 · 9 citations
- Perturbation CheckLists for Evaluating NLG Evaluation MetricsAnanya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan et al.EMNLP 2021 · 32 citations
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai et al.ACL 2024
