BLESS: Benchmarking Large Language Models on Sentence Simplification
Tannon Kew, Alison Chi, Laura Vásquez-Rodríguez, Sweta Agrawal, Dennis Aumiller, Fernando Alva-Manchego, Matthew Shardlow
Abstract
We present BLESS, a comprehensive performance benchmark of the most recent state-ofthe-art large language models (LLMs) on the task of text simplification (TS). We examine how well off-the-shelf LLMs can solve this challenging task, assessing a total of 44 models, differing in size, architecture, pre-training methods, and accessibility, on three test sets from different domains (Wikipedia, news, and medical) under a few-shot setting. Our analysis considers a suite of automatic metrics as well as a large-scale quantitative investigation into the types of common edit operations performed by the different models. Furthermore, we perform a manual qualitative analysis on a subset of model outputs to better gauge the quality of the generated simplifications. Our evaluation indicates that the best LLMs, despite not being trained on TS, perform comparably with state-of-the-art TS baselines. Additionally, we find that certain LLMs demonstrate a greater range and diversity of edit operations. Our performance benchmark will be available as a resource for the development of future TS methods and evaluation metrics. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7095a84-d0fd-4d90-aabb-a3dc8613fb8bCited by top-tier papers6
- Digital Comprehensibility Assessment of Simplified Texts among Persons with Intellectual DisabilitiesAndreas Säuberli, Franz Holzknecht, Patrick Haller, Silvana Deilen et al.CHI 2024 · 10 citations
- Evaluating LLMs for Targeted Concept Simplification for Domain-Specific TextsSumit Asthana, Hannah Rashkin, Elizabeth Clark, Fantine Huot et al.EMNLP 2024 · 2 citations
- Attacking Misinformation Detection Using Adversarial Examples Generated by Language ModelsPiotr Przybyla, Euan McGill, Horacio SaggionEMNLP 2025 · 1 citation
- JUDGEBERT: Assessing Legal Meaning Preservation Between SentencesDavid Beauchemin, Michelle Albert-Rochette, Richard Khoury, Pierre-Luc DézielEMNLP 2025
- QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability JudgmentsDavid Beauchemin, Richard KhouryEMNLP 2025
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
Related papers
- Evaluating LLMs for Portuguese Sentence Simplification with Linguistic InsightsArthur Mariano Rocha De Azevedo Scalercio, Elvis A. de Souza, Maria José Bocorny Finatto, Aline PaesACL 2025 · 2 citations
- Revisiting non-English Text Simplification: A Unified Multilingual BenchmarkMichael J. Ryan, Tarek Naous, Wei XuACL 2023 · 14 citations
- LENS: A Learnable Evaluation Metric for Text SimplificationMounica Maddela, Yao Dou, David Heineman, Wei XuACL 2023 · 21 citations
- Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSADavid Heineman, Yao Dou, Mounica Maddela, Wei XuEMNLP 2023 · 5 citations
- XCodeEval: An Execution-based Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and RetrievalMohammad Abdullah Matin Khan, M. Saiful Bari, Xuan Do Long, Weishi Wang et al.ACL 2024 · 21 citations
