GENIE: Toward Reproducible and Standardized Human Evaluation for Text Generation
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, Daniel S. Weld
摘要
While often assumed a gold standard, effective human evaluation of text generation remains an important, open area for research. We revisit this problem with a focus on producing consistent evaluations that are reproducible-over time and across different populations. We study this goal in different stages of the human evaluation pipeline. In particular, we consider design choices for the annotation interface used to elicit human judgments and their impact on reproducibility. Furthermore, we develop an automated mechanism for maintaining annotator quality via a probabilistic model that detects and excludes noisy annotators. Putting these lessons together, we introduce GENIE: a system for running standardized human evaluations across different generation tasks. We instantiate GENIE with datasets representing four core challenges in text generation: machine translation, summarization, commonsense reasoning, and machine comprehension. For each task, GENIE offers a leaderboard that automatically crowdsources annotations for submissions, evaluating them along axes such as correctness, conciseness, and fluency. We have made the GE-NIE leaderboards publicly available, and have already ranked 50 submissions from 10 different research groups. 1 We hope GENIE encourages further progress toward effective, standardized evaluations for text generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang 等ICCV 2023 · 被引用 400 次
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim 等CHI 2024 · 被引用 81 次
- Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchySimon Ging, María Alejandra Bravo, Thomas BroxICLR 2024 · 被引用 24 次
- Foundational Autoraters: Taming Large Language Models for Better Automatic EvaluationTu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar 等EMNLP 2024 · 被引用 14 次
- Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement TheoryZiang Xiao, Susu Zhang, Vivian Lai, Q. Vera LiaoEMNLP 2023 · 被引用 6 次
它引用的顶会 Paper10
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi 等ICLR 2020 · 被引用 521 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- Back to the Future: Unsupervised Backprop-based Decoding for Counterfactual and Abductive Commonsense ReasoningLianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula 等EMNLP 2020 · 被引用 25 次
相关 Paper
- Toward Verifiable and Reproducible Human Evaluation for Text-to-Image GenerationMayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami 等CVPR 2023
- Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text GenerationKatelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi, Zongwan Cao 等ACL 2026
- QGEval: Benchmarking Multi-dimensional Evaluation for Question GenerationWeiping Fu, Bifan Wei, Jianxiang Hu, Zhongmin Cai 等EMNLP 2024 · 被引用 8 次
- Perturbation CheckLists for Evaluating NLG Evaluation MetricsAnanya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan 等EMNLP 2021 · 被引用 32 次
- Re-evaluating Evaluation in Text SummarizationManik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu 等EMNLP 2020 · 被引用 3 次
