Perturbation CheckLists for Evaluating NLG Evaluation Metrics
Ananya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, Mitesh M. Khapra
摘要
Natural Language Generation (NLG) evaluation is a multifaceted task requiring assessment of multiple desirable criteria, e.g., fluency, coherency, coverage, relevance, adequacy, overall quality, etc. Across existing datasets for 6 NLG tasks, we observe that the human evaluation scores on these multiple criteria are often not correlated. For example, there is a very low correlation between human scores on fluency and data coverage for the task of structured data to text generation. This suggests that the current recipe of proposing new automatic evaluation metrics for NLG by showing that they correlate well with scores assigned by humans for a single criteria (overall quality) alone is inadequate. Indeed, our extensive study involving 25 automatic evaluation metrics across 6 different tasks and 18 different evaluation criteria shows that there is no single metric which correlates well with human scores on all desirable criteria, for most NLG tasks. Given this situation, we propose CheckLists for better design and evaluation of automatic metrics. We design templates which target a specific criteria (e.g., coverage) and perturb the output such that the quality gets affected only along this specific criteria (e.g., the coverage drops). We show that existing evaluation metrics are not robust against even such simple perturbations and disagree with scores assigned by humans to the perturbed output. The proposed templates thus allow for a fine-grained assessment of automatic evaluation metrics exposing their limitations and will facilitate better design, analysis and evaluation of such metrics. 1 Task Criteria Machine Translation Adequacy: The generated translation should adequately represent all the information present in the reference. Question Generation Relevance: Is the question related to the source material they are based upon. Answerability: Is the generated question answerable given the context. Informativeness: The summary should convey the key points of the text. Non-redundancy: The summary should not repeat any points, and ideally have maximal information coverage within the limited text length. Abstractive Summarization Referential clarity: Any intra-sentence or cross-sentence references in the summary should be unambiguous and within the scope of the summary. Focus: The summary needs to have a focus and all the sentences need to contain information related to this focal point. Structure and Coherence: The summary should be a well-organized and coherent body of information Dialogue Generation Making sense: Does the bot say things that don't make sense? Engagingness: Is the dialogue agent enjoyable to talk to? Interestingness: Did you find the bot interesting to talk to? Inquisitivenes: Does the bot ask a good amount of questions? Listening: Does the bot pay attention to what you say? Avoiding Repetition: Does the bot repeat itself? (either within or across utterances) Humanness: Is the conversation with a person or a bot? Image Captioning Relevance: The caption should be specific and related to the image. Thoroughness: The caption should adequately describe the image. Data Coverage: Does the text include descriptions of all predicates presented in the data? Relevance: Does the text describe only such predicates which are found in the data?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsChen Zhang, Luis Fernando D'Haro, Yiming Chen, Malu Zhang 等AAAI 2024 · 被引用 57 次
- DEMETR: Diagnosing Evaluation Metrics for TranslationMarzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song 等EMNLP 2022 · 被引用 18 次
- FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationChen Zhang, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs 等EMNLP 2022 · 被引用 13 次
- NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference ChecklistIftitahu Ni'mah, Meng Fang, Vlado Menkovski, Mykola PechenizkiyACL 2023 · 被引用 9 次
- IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian LanguagesAnanya B. Sai, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan 等ACL 2023 · 被引用 8 次
它引用的顶会 Paper3
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 被引用 51 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation MetricsNitika Mathur, Timothy Baldwin, Trevor CohnACL 2020 · 被引用 14 次
相关 Paper
- Compression, Transduction, and Creation: A Unified Framework for Evaluating Natural Language GenerationMingkai Deng, Bowen Tan, Zhengzhong Liu, Eric P. Xing 等EMNLP 2021 · 被引用 49 次
- Towards a Unified Multi-Dimensional Evaluator for Text GenerationMing Zhong, Yang Liu, Da Yin, Yuning Mao 等EMNLP 2022 · 被引用 103 次
- Spurious Correlations in Reference-Free Evaluation of Text GenerationEsin Durmus, Faisal Ladhak, Tatsunori HashimotoACL 2022
- Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement TheoryZiang Xiao, Susu Zhang, Vivian Lai, Q. Vera LiaoEMNLP 2023 · 被引用 6 次
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu 等ACL 2021
