Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSA
David Heineman, Yao Dou, Mounica Maddela, Wei Xu
Abstract
Large language models (e.g., GPT-4) are uniquely capable of producing highly rated text simplification, yet current human evaluation methods fail to provide a clear understanding of systems' specific strengths and weaknesses. To address this limitation, we introduce SALSA, an edit-based human annotation framework that enables holistic and fine-grained text simplification evaluation. We develop twenty one linguistically grounded edit types, covering the full spectrum of success and failure across dimensions of conceptual, syntactic and lexical simplicity. Using SALSA, we collect 19K edit annotations on 840 simplifications, revealing discrepancies in the distribution of simplification strategies performed by fine-tuned models, prompted LLMs and humans, and find GPT-3.5 performs more quality edits than humans, but still exhibits frequent errors. Using our finegrained annotations, we develop LENS-SALSA, a reference-free automatic simplification metric, trained to predict sentence-and word-level quality simultaneously. Additionally, we introduce word-level quality estimation for simplification and report promising baseline results. Our data, new metric, and annotation toolkit are available at https://salsa-eval.com . EXAMPLE Zero-shot GPT-3.5 On 14 November, an interview with journalist Piers Morgan was published, where Ronaldo said ... On 14 November, Piers Morgan interviewed Ronaldo, who expressed ...
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cf39c7e-67e3-4406-b1a1-a2b1031239afCited by top-tier papers5
- Reducing Privacy Risks in Online Self-Disclosures with Language ModelsYao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra et al.ACL 2024 · 14 citations
- AutoMetrics: Approximate Human Judgments with Automatically Generated EvaluatorsMichael J Ryan, Yanzhe Zhang, Amol Salunkhe, Yi Chu et al.ICLR 2026 · 3 citations
- Evaluating LLMs for Portuguese Sentence Simplification with Linguistic InsightsArthur Mariano Rocha De Azevedo Scalercio, Elvis A. de Souza, Maria José Bocorny Finatto, Aline PaesACL 2025 · 2 citations
- Improving Minimum Bayes Risk Decoding with Multi-PromptDavid Heineman, Yao Dou, Wei XuEMNLP 2024 · 1 citation
- Label Confidence Weighted Learning for Target-level Sentence SimplificationXin Ying Qiu, Jingshen ZhangEMNLP 2024 · 1 citation
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri et al.NeurIPS 2023 · 516 citations
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong et al.ACL 2020 · 103 citations
Related papers
- LENS: A Learnable Evaluation Metric for Text SimplificationMounica Maddela, Yao Dou, David Heineman, Wei XuACL 2023 · 21 citations
- BLESS: Benchmarking Large Language Models on Sentence SimplificationTannon Kew, Alison Chi, Laura Vásquez-Rodríguez, Sweta Agrawal et al.EMNLP 2023 · 15 citations
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai et al.ACL 2024
- Linguistic Corpus Annotation for Automatic Text Simplification EvaluationRémi Cardon, Adrien Bibal, Rodrigo Wilkens, David Alfter et al.EMNLP 2022 · 5 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
