On the Blind Spots of Model-Based Evaluation Metrics for Text Generation
Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James R. Glass, Yulia Tsvetkov
Abstract
In this work, we explore a useful but often neglected methodology for robustness analysis of text generation evaluation metrics: stress tests with synthetic data. Basically, we design and synthesize a wide range of potential errors and check whether they result in a commensurate drop in the metric scores. We examine a range of recently proposed evaluation metrics based on pretrained language models, for the tasks of open-ended generation, translation, and summarization. Our experiments reveal interesting insensitivities, biases, or even loopholes in existing metrics. For example, we find that BERTScore is confused by truncation errors in summarization, and MAUVE (built on top of GPT-2) is insensitive to errors at the beginning or middle of generations. Further, we investigate the reasons behind these blind spots and suggest practical workarounds for a more reliable evaluation of text generation. We have released our code and data at https://github. com/cloudygoose/blindspot_nlg .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a52527ab-c721-41af-8602-181dbf6c2ce8Cited by top-tier papers13
- Understanding In-Context Learning via Supportive Pretraining DataXiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov et al.ACL 2023 · 16 citations
- BLESS: Benchmarking Large Language Models on Sentence SimplificationTannon Kew, Alison Chi, Laura Vásquez-Rodríguez, Sweta Agrawal et al.EMNLP 2023 · 15 citations
- CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model GenerationPei Ke, Bosi Wen, Andrew Feng, Xiao Liu et al.ACL 2024 · 9 citations
- SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated ResponsesDongwei Jiang, Jingyu Zhang, Orion Weller, Nathaniel Weir et al.AAAI 2025 · 8 citations
- APPLS: Evaluating Evaluation Metrics for Plain Language SummarizationYue Guo, Tal August, Gondy Leroy, Trevor Cohen et al.EMNLP 2024 · 7 citations
Builds on32
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
Related papers
- Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic FactorsMarvin Kaster, Wei Zhao, Steffen EgerEMNLP 2021 · 13 citations
- Reproducibility Issues for BERT-based Evaluation MetricsYanran Chen, Jonas Belouadi, Steffen EgerEMNLP 2022 · 11 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- FrugalScore: Learning Cheaper, Lighter and Faster Evaluation Metrics for Automatic Text GenerationMoussa Kamal Eddine, Guokan Shang, Antoine J.-P. Tixier, Michalis VazirgiannisACL 2022
