Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSA
David Heineman, Yao Dou, Mounica Maddela, Wei Xu
摘要
Large language models (e.g., GPT-4) are uniquely capable of producing highly rated text simplification, yet current human evaluation methods fail to provide a clear understanding of systems' specific strengths and weaknesses. To address this limitation, we introduce SALSA, an edit-based human annotation framework that enables holistic and fine-grained text simplification evaluation. We develop twenty one linguistically grounded edit types, covering the full spectrum of success and failure across dimensions of conceptual, syntactic and lexical simplicity. Using SALSA, we collect 19K edit annotations on 840 simplifications, revealing discrepancies in the distribution of simplification strategies performed by fine-tuned models, prompted LLMs and humans, and find GPT-3.5 performs more quality edits than humans, but still exhibits frequent errors. Using our finegrained annotations, we develop LENS-SALSA, a reference-free automatic simplification metric, trained to predict sentence-and word-level quality simultaneously. Additionally, we introduce word-level quality estimation for simplification and report promising baseline results. Our data, new metric, and annotation toolkit are available at https://salsa-eval.com . EXAMPLE Zero-shot GPT-3.5 On 14 November, an interview with journalist Piers Morgan was published, where Ronaldo said ... On 14 November, Piers Morgan interviewed Ronaldo, who expressed ...
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Reducing Privacy Risks in Online Self-Disclosures with Language ModelsYao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra 等ACL 2024 · 被引用 14 次
- AutoMetrics: Approximate Human Judgments with Automatically Generated EvaluatorsMichael J Ryan, Yanzhe Zhang, Amol Salunkhe, Yi Chu 等ICLR 2026 · 被引用 3 次
- Evaluating LLMs for Portuguese Sentence Simplification with Linguistic InsightsArthur Mariano Rocha De Azevedo Scalercio, Elvis A. de Souza, Maria José Bocorny Finatto, Aline PaesACL 2025 · 被引用 2 次
- Improving Minimum Bayes Risk Decoding with Multi-PromptDavid Heineman, Yao Dou, Wei XuEMNLP 2024 · 被引用 1 次
- Label Confidence Weighted Learning for Target-level Sentence SimplificationXin Ying Qiu, Jingshen ZhangEMNLP 2024 · 被引用 1 次
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri 等NeurIPS 2023 · 被引用 516 次
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong 等ACL 2020 · 被引用 103 次
相关 Paper
- LENS: A Learnable Evaluation Metric for Text SimplificationMounica Maddela, Yao Dou, David Heineman, Wei XuACL 2023 · 被引用 21 次
- BLESS: Benchmarking Large Language Models on Sentence SimplificationTannon Kew, Alison Chi, Laura Vásquez-Rodríguez, Sweta Agrawal 等EMNLP 2023 · 被引用 15 次
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai 等ACL 2024
- Linguistic Corpus Annotation for Automatic Text Simplification EvaluationRémi Cardon, Adrien Bibal, Rodrigo Wilkens, David Alfter 等EMNLP 2022 · 被引用 5 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
