Evaluating the Evaluators: Are readability metrics good measures of readability?
Isabel Cachola, Daniel Khashabi, Mark Dredze
Abstract
Plain Language Summarization (PLS) aims to distill complex documents into accessible summaries for non-expert audiences. In this paper, we conduct a thorough survey of PLS literature, and identify that the current standard practice for readability evaluation is to use traditional readability metrics, such as Flesch-Kincaid Grade Level (FKGL). However, despite proven utility in other fields, these metrics have not been compared to human readability judgments in PLS. We evaluate 8 readability metrics and show that most correlate poorly with human judgments, including the most popular metric, FKGL. We then show that Language Models (LMs) are better judges of readability, with the best-performing model achieving a Pearson correlation of 0.56 with human judgments. Extending our analysis to PLS datasets, which contain summaries aimed at non-expert audiences, we find that LMs better capture deeper measures of readability, such as required background knowledge, and lead to different conclusions than the traditional metrics. Based on these findings, we offer recommendations for best practices in the evaluation of plain language summaries. We release our analysis code and survey data. JHU-CLSP/eval-the-eval-readability William H DuBay. 2004. The principles of readability. Impact Information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7d617ce-dac3-4575-8b0f-64d850f0d74bCited by top-tier papers1
Ask how each one uses itBuilds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Automated Lay Language Summarization of Biomedical Scientific ReviewsYue Guo, Wei Qiu, Yizhong Wang, Trevor CohenAAAI 2021 · 100 citations
- Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationYixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao et al.ACL 2023 · 50 citations
Related papers
- Know Your Audience: The benefits and pitfalls of generating plain language summaries beyond the "general" audienceTal August, Kyle Lo, Noah A. Smith, Katharina ReineckeCHI 2024 · 11 citations
- APPLS: Evaluating Evaluation Metrics for Plain Language SummarizationYue Guo, Tal August, Gondy Leroy, Trevor Cohen et al.EMNLP 2024 · 7 citations
- No Reader Left Behind: Multi-Agent Summaries Everyone Can UnderstandJimin Jung, MyoungJin Kim, Jaehyung Seo, Heuiseok LimACL 2026
- Play the Shannon Game with Language Models: A Human-Free Approach to Summary EvaluationNicholas Egan, Oleg V. Vasilyev, John BohannonAAAI 2022 · 22 citations
- DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question AnsweringPei Ke, Fei Huang, Fei Mi, Yasheng Wang et al.ACL 2023 · 2 citations
