Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
Katelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi, Zongwan Cao, Chenjun Xu, Bingbing Wen, Su Lin Blodgett, Lucy Lu Wang
Abstract
Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and welldocumented protocols-details that are frequently missing in current practice. In this work, we conduct a large-scale analysis of human evaluation protocols for evaluating longform generation tasks in *CL conference publications from 2023-2025, including a full manual review of 284 papers and LLM-assisted analysis for another 1.8k+ papers. We define a set of 20 reportable criteria related to reproducibility of human evaluation studies, and apply these criteria to systematically examine reporting norms and practices within the community. We find widespread under-reporting of important aspects of human evaluation study design, leading to ambiguity about what was measured and how, who contributed judgments, and how judgments should be interpreted. Based on these findings, we outline actionable recommendations to support more transparent and reproducible reporting in future research. 1 * Denotes equal contribution 1 Our analysis code and annotated dataset can be found at: https://github.com/larchlab/Illusions-of-the-Gold-Standard .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Impact of Annotator Demographics on Sentiment Dataset LabelingYi Ding, Jacob You, Tonja-Katrin Machulla, Jennifer Jacobs et al.CSCW 2022 · 20 citations
- ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language ModelsAparna Elangovan, Ling Liu, Lei Xu, Sravan Babu Bodapati et al.ACL 2024 · 19 citations
- Automated Metrics for Medical Multi-Document Summarization Disagree with Human EvaluationsLucy Lu Wang, Yulia Otmakhova, Jay DeYoung, Thinh Hung Truong et al.ACL 2023 · 12 citations
Related papers
- Toward Verifiable and Reproducible Human Evaluation for Text-to-Image GenerationMayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami et al.CVPR 2023
- LLMEval: A Preliminary Study on How to Evaluate Large Language ModelsYue Zhang, Ming Zhang, Haipeng Yuan, Shichun Liu et al.AAAI 2024 · 29 citations
- GENIE: Toward Reproducible and Standardized Human Evaluation for Text GenerationDaniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie et al.EMNLP 2022 · 13 citations
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman et al.EMNLP 2024 · 47 citations
- L-Eval: Instituting Standardized Evaluation for Long Context Language ModelsChenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao et al.ACL 2024 · 6 citations
