Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement Theory
Ziang Xiao, Susu Zhang, Vivian Lai, Q. Vera Liao
摘要
We address a fundamental challenge in Natural Language Generation (NLG) model evaluation-the design and evaluation of evaluation metrics. Recognizing the limitations of existing automatic metrics and noises from how current human evaluation was conducted, we propose METRICEVAL, a framework informed by measurement theory, the foundation of educational test design, for conceptualizing and evaluating the reliability and validity of NLG evaluation metrics. The framework formalizes the source of measurement error and offers statistical tools for evaluating evaluation metrics based on empirical data. With our framework, one can quantify the uncertainty of the metrics to better interpret the result. To exemplify the use of our framework in practice, we analyzed a set of evaluation metrics for summarization and identified issues related to conflated validity structure in human-eval and reliability in LLM-based metrics. Through METRICEVAL 1 , we aim to promote the design, evaluation, and interpretation of valid and reliable metrics to advance robust and effective NLG models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme 等ACL 2024 · 被引用 27 次
- A Comparative Analysis of Information Gathering by Chatbots, Questionnaires, and Humans in Clinical Pre-ConsultationBrenna Li, Saba Tauseef, Khai N. Truong, Alex MariakakisCHI 2025 · 被引用 16 次
- AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value DifferenceJing Yao, Shitong Duan, Xiaoyuan Yi, Dongkuan Xu 等ICLR 2026 · 被引用 4 次
- ECBD: Evidence-Centered Benchmark Design for NLPYu Lu Liu, Su Lin Blodgett, Jackie C. K. Cheung, Vera Liao 等ACL 2024 · 被引用 3 次
- Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the WildWillem van der Maden, Malak Sadek, Ziang Xiao, Aske Mottelson 等CHI 2026 · 被引用 2 次
它引用的顶会 Paper11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
相关 Paper
- Towards a Unified Multi-Dimensional Evaluator for Text GenerationMing Zhong, Yang Liu, Da Yin, Yuning Mao 等EMNLP 2022 · 被引用 103 次
- DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question AnsweringPei Ke, Fei Huang, Fei Mi, Yasheng Wang 等ACL 2023 · 被引用 2 次
- Perturbation CheckLists for Evaluating NLG Evaluation MetricsAnanya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan 等EMNLP 2021 · 被引用 32 次
- Spurious Correlations in Reference-Free Evaluation of Text GenerationEsin Durmus, Faisal Ladhak, Tatsunori HashimotoACL 2022
- A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better InterpretabilityXinyu Hu, Mingqi Gao, Li Lin, Zhenghan Yu 等ACL 2025
