Perturbation CheckLists for Evaluating NLG Evaluation Metrics
Ananya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, Mitesh M. Khapra
Abstract
Natural Language Generation (NLG) evaluation is a multifaceted task requiring assessment of multiple desirable criteria, e.g., fluency, coherency, coverage, relevance, adequacy, overall quality, etc. Across existing datasets for 6 NLG tasks, we observe that the human evaluation scores on these multiple criteria are often not correlated. For example, there is a very low correlation between human scores on fluency and data coverage for the task of structured data to text generation. This suggests that the current recipe of proposing new automatic evaluation metrics for NLG by showing that they correlate well with scores assigned by humans for a single criteria (overall quality) alone is inadequate. Indeed, our extensive study involving 25 automatic evaluation metrics across 6 different tasks and 18 different evaluation criteria shows that there is no single metric which correlates well with human scores on all desirable criteria, for most NLG tasks. Given this situation, we propose CheckLists for better design and evaluation of automatic metrics. We design templates which target a specific criteria (e.g., coverage) and perturb the output such that the quality gets affected only along this specific criteria (e.g., the coverage drops). We show that existing evaluation metrics are not robust against even such simple perturbations and disagree with scores assigned by humans to the perturbed output. The proposed templates thus allow for a fine-grained assessment of automatic evaluation metrics exposing their limitations and will facilitate better design, analysis and evaluation of such metrics. 1 Task Criteria Machine Translation Adequacy: The generated translation should adequately represent all the information present in the reference. Question Generation Relevance: Is the question related to the source material they are based upon. Answerability: Is the generated question answerable given the context. Informativeness: The summary should convey the key points of the text. Non-redundancy: The summary should not repeat any points, and ideally have maximal information coverage within the limited text length. Abstractive Summarization Referential clarity: Any intra-sentence or cross-sentence references in the summary should be unambiguous and within the scope of the summary. Focus: The summary needs to have a focus and all the sentences need to contain information related to this focal point. Structure and Coherence: The summary should be a well-organized and coherent body of information Dialogue Generation Making sense: Does the bot say things that don't make sense? Engagingness: Is the dialogue agent enjoyable to talk to? Interestingness: Did you find the bot interesting to talk to? Inquisitivenes: Does the bot ask a good amount of questions? Listening: Does the bot pay attention to what you say? Avoiding Repetition: Does the bot repeat itself? (either within or across utterances) Humanness: Is the conversation with a person or a bot? Image Captioning Relevance: The caption should be specific and related to the image. Thoroughness: The caption should adequately describe the image. Data Coverage: Does the text include descriptions of all predicates presented in the data? Relevance: Does the text describe only such predicates which are found in the data?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93f65454-b68f-4185-b43f-20a21f3acefcCited by top-tier papers12
- A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsChen Zhang, Luis Fernando D'Haro, Yiming Chen, Malu Zhang et al.AAAI 2024 · 57 citations
- DEMETR: Diagnosing Evaluation Metrics for TranslationMarzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song et al.EMNLP 2022 · 18 citations
- FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationChen Zhang, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs et al.EMNLP 2022 · 13 citations
- NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference ChecklistIftitahu Ni'mah, Meng Fang, Vlado Menkovski, Mykola PechenizkiyACL 2023 · 9 citations
- IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian LanguagesAnanya B. Sai, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan et al.ACL 2023 · 8 citations
Builds on3
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation MetricsNitika Mathur, Timothy Baldwin, Trevor CohnACL 2020 · 14 citations
Related papers
- Compression, Transduction, and Creation: A Unified Framework for Evaluating Natural Language GenerationMingkai Deng, Bowen Tan, Zhengzhong Liu, Eric P. Xing et al.EMNLP 2021 · 49 citations
- Towards a Unified Multi-Dimensional Evaluator for Text GenerationMing Zhong, Yang Liu, Da Yin, Yuning Mao et al.EMNLP 2022 · 103 citations
- Spurious Correlations in Reference-Free Evaluation of Text GenerationEsin Durmus, Faisal Ladhak, Tatsunori HashimotoACL 2022
- Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement TheoryZiang Xiao, Susu Zhang, Vivian Lai, Q. Vera LiaoEMNLP 2023 · 6 citations
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu et al.ACL 2021
