LENS: A Learnable Evaluation Metric for Text Simplification
Mounica Maddela, Yao Dou, David Heineman, Wei Xu
摘要
Training learnable metrics using modern language models has recently emerged as a promising method for the automatic evaluation of machine translation. However, existing human evaluation datasets for text simplification have limited annotations that are based on unitary or outdated models, making them unsuitable for this approach. To address these issues, we introduce the SIMPEVAL corpus that contains: SIMPEVAL PAST , comprising 12K human ratings on 2.4K simplifications of 24 past systems, and SIMPEVAL 2022 , a challenging simplification benchmark consisting of over 1K human ratings of 360 simplifications including GPT-3.5 generated text. Training on SIMPEVAL, we present LENS, a Learnable Evaluation Metric for Text Simplification. Extensive empirical results show that LENS correlates much better with human judgment than existing metrics, paving the way for future progress in the evaluation of text simplification. We also introduce RANK & RATE, a human evaluation framework that rates simplifications from several models in a list-wise manner using an interactive interface, which ensures both consistency and accuracy in the evaluation process and is used to create the SIMPEVAL datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Beyond the Chat: Executable and Verifiable Text-Editing with LLMsPhilippe Laban, Jesse Vig, Marti A. Hearst, Caiming Xiong 等UIST 2024 · 被引用 29 次
- Polos: Multimodal Metric Learning from Human Feedback for Image CaptioningYuiga Wada, Kanta Kaneda, Daichi Saito, Komei SugiuraCVPR 2024 · 被引用 16 次
- BLESS: Benchmarking Large Language Models on Sentence SimplificationTannon Kew, Alison Chi, Laura Vásquez-Rodríguez, Sweta Agrawal 等EMNLP 2023 · 被引用 15 次
- Foundational Autoraters: Taming Large Language Models for Better Automatic EvaluationTu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar 等EMNLP 2024 · 被引用 14 次
- Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through EditsTuhin Chakrabarty, Philippe Laban, Chien-Sheng WuCHI 2025 · 被引用 14 次
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong 等ACL 2020 · 被引用 103 次
相关 Paper
- Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSADavid Heineman, Yao Dou, Mounica Maddela, Wei XuEMNLP 2023 · 被引用 5 次
- Linguistic Corpus Annotation for Automatic Text Simplification EvaluationRémi Cardon, Adrien Bibal, Rodrigo Wilkens, David Alfter 等EMNLP 2022 · 被引用 5 次
- Evaluating LLMs for Portuguese Sentence Simplification with Linguistic InsightsArthur Mariano Rocha De Azevedo Scalercio, Elvis A. de Souza, Maria José Bocorny Finatto, Aline PaesACL 2025 · 被引用 2 次
- Revisiting non-English Text Simplification: A Unified Multilingual BenchmarkMichael J. Ryan, Tarek Naous, Wei XuACL 2023 · 被引用 14 次
- SIMSUM: Document-level Text Simplification via Simultaneous SummarizationSofia Blinova, Xinyu Zhou, Martin Jaggi, Carsten Eickhoff 等ACL 2023 · 被引用 11 次
