The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation
Marzena Karpinska, Nader Akoury, Mohit Iyyer
Abstract
Recent text generation research has increasingly focused on open-ended domains such as story and poetry generation. Because models built for such tasks are difficult to evaluate automatically, most researchers in the space justify their modeling choices by collecting crowdsourced human judgments of text quality (e.g., Likert scores of coherence or grammaticality) from Amazon Mechanical Turk (AMT). In this paper, we first conduct a survey of 45 open-ended text generation papers and find that the vast majority of them fail to report crucial details about their AMT tasks, hindering reproducibility. We then run a series of story evaluation experiments with both AMT workers and English teachers and discover that even with strict qualification filters, AMT workers (unlike teachers) fail to distinguish between model-generated text and human-generated references. We show that AMT worker judgments improve when they are shown model-generated output alongside human-generated references, which enables the workers to better calibrate their ratings. Finally, interviews with the English teachers provide deeper insights into the challenges of the evaluation process, particularly when rating model-generated text.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers39
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu et al.ICLR 2024 · 871 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng et al.ICLR 2024 · 299 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
Builds on19
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- MEGATRON-CNTRL: Controllable Story Generation with External Knowledge Using Large-Scale Language ModelsPeng Xu, Mostofa Patwary, Mohammad Shoeybi, Raul Puri et al.EMNLP 2020 · 104 citations
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong et al.ACL 2020 · 103 citations
- PlotMachines: Outline-Conditioned Generation with Dynamic Plot State TrackingHannah Rashkin, Asli Celikyilmaz, Yejin Choi, Jianfeng GaoEMNLP 2020 · 100 citations
- With Little Power Comes Great ResponsibilityDallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia et al.EMNLP 2020 · 76 citations
Related papers
- Incorporating Worker Perspectives into MTurk Annotation Practices for NLPOlivia Huang, Eve Fleisig, Dan KleinEMNLP 2023 · 1 citation
- A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for SummarizationLining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch et al.ACL 2023 · 6 citations
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu et al.ACL 2021
- Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text GenerationKatelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi, Zongwan Cao et al.ACL 2026
- GENIE: Toward Reproducible and Standardized Human Evaluation for Text GenerationDaniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie et al.EMNLP 2022 · 13 citations
