Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation
Francesco Moramarco, Alex Papadopoulos-Korfiatis, Mark Perera, Damir Juric, Jack Flann, Ehud Reiter, Anya Belz, Aleksandar Savkov
Abstract
In recent years, machine learning models have rapidly become better at generating clinical consultation notes; yet, there is little work on how to properly evaluate the generated consultation notes to understand the impact they may have on both the clinician using them and the patient's clinical safety. To address this we present an extensive human evaluation study of consultation notes where 5 clinicians (i) listen to 57 mock consultations, (ii) write their own notes, (iii) post-edit a number of automatically generated notes, and (iv) extract all the errors, both quantitative and qualitative. We then carry out a correlation study with 18 automatic quality metrics and the human judgements. We find that a simple, character-based Levenshtein distance metric performs on par if not better than common model-based metrics like BertScore. All our findings and annotations are open-sourced. Related Work Note Generation has been in the focus of the academic community with both extractive methods (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00a9728c-c200-4891-9923-551c31475a58Cited by top-tier papers1
Ask how each one uses itBuilds on4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology ReportsYuhao Zhang, Derek Merck, Emily Bao Tsai, Christopher D. Manning et al.ACL 2020 · 160 citations
- Generating SOAP Notes from Doctor-Patient Conversations Using Modular Summarization TechniquesKundan Krishna, Sopan Khosla, Jeffrey P. Bigham, Zachary C. LiptonACL 2021
Related papers
- Identifying Reliable Evaluation Metrics for Scientific Text RevisionLéane Jourdan, Nicolas Hernandez, Florian Boudin, Richard DufourACL 2025
- DocLens: Multi-aspect Fine-grained Medical Text EvaluationYiqing Xie, Sheng Zhang, Hao Cheng, Pengfei Liu et al.ACL 2024
- CREPE: Rapid Chest X-ray Report Evaluation by Predicting Multi-category Error CountsGihun Cho, Seunghyun Jang, Hanbin Ko, Inhyeok Baek et al.EMNLP 2025
- Language Model Augmented Relevance ScoreRuibo Liu, Jason Wei, Soroush VosoughiACL 2021
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation ModelsWangchunshu Zhou, Ke XuAAAI 2020 · 49 citations
