APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
Yue Guo, Tal August, Gondy Leroy, Trevor Cohen, Lucy Lu Wang
Abstract
While there has been significant development of models for Plain Language Summarization (PLS), evaluation remains a challenge. PLS lacks a dedicated assessment metric, and the suitability of text generation evaluation metrics is unclear due to the unique transformations involved (e.g., adding background explanations, removing jargon). To address these questions, our study introduces a granular meta-evaluation testbed, APPLS, designed to evaluate metrics for PLS. We identify four PLS criteria from previous work-informativeness, simplification, coherence, and faithfulness-and define a set of perturbations corresponding to these criteria that sensitive metrics should be able to detect. We apply these perturbations to the texts of two PLS datasets to create our testbed. Using AP-PLS, we assess performance of 14 metrics, including automated scores, lexical features, and LLM prompt-based evaluations. Our analysis reveals that while some current metrics show sensitivity to specific criteria, no single method captures all four criteria simultaneously. We therefore recommend a suite of automated metrics be used to capture PLS quality along all relevant criteria. This work contributes the first meta-evaluation testbed for PLS and a comprehensive evaluation of existing metrics. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Evaluating the Evaluators: Are readability metrics good measures of readability?Isabel Cachola, Daniel Khashabi, Mark DredzeEMNLP 2025 · 1 citation
- Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian LanguagesAmir Hossein Yari, Kalmit Kulkarni, Ahmad Raza Khan, Fajri KotoACL 2026
- Evaluating Taxonomy Free Character Role Labeling (TF-CRL) in News Stories using Large Language ModelsDavid G. Hobson, Derek Ruths, Andrew PiperEMNLP 2025
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Automated Lay Language Summarization of Biomedical Scientific ReviewsYue Guo, Wei Qiu, Yizhong Wang, Trevor CohenAAAI 2021 · 100 citations
- Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsSaku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko AizawaAAAI 2020 · 92 citations
- Perturbation CheckLists for Evaluating NLG Evaluation MetricsAnanya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan et al.EMNLP 2021 · 32 citations
Related papers
- DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question AnsweringPei Ke, Fei Huang, Fei Mi, Yasheng Wang et al.ACL 2023 · 2 citations
- Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language ModelQi Jia, Siyu Ren, Yizhu Liu, Kenny Q. ZhuEMNLP 2023 · 4 citations
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai et al.ACL 2024
- PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization EvaluationChristoph Leiter, Steffen EgerEMNLP 2024 · 5 citations
- Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement TheoryZiang Xiao, Susu Zhang, Vivian Lai, Q. Vera LiaoEMNLP 2023 · 6 citations
