The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation
Marzena Karpinska, Nader Akoury, Mohit Iyyer
摘要
Recent text generation research has increasingly focused on open-ended domains such as story and poetry generation. Because models built for such tasks are difficult to evaluate automatically, most researchers in the space justify their modeling choices by collecting crowdsourced human judgments of text quality (e.g., Likert scores of coherence or grammaticality) from Amazon Mechanical Turk (AMT). In this paper, we first conduct a survey of 45 open-ended text generation papers and find that the vast majority of them fail to report crucial details about their AMT tasks, hindering reproducibility. We then run a series of story evaluation experiments with both AMT workers and English teachers and discover that even with strict qualification filters, AMT workers (unlike teachers) fail to distinguish between model-generated text and human-generated references. We show that AMT worker judgments improve when they are shown model-generated output alongside human-generated references, which enables the workers to better calibrate their ratings. Finally, interviews with the English teachers provide deeper insights into the challenges of the evaluation process, particularly when rating model-generated text.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu 等ICLR 2024 · 被引用 871 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng 等ICLR 2024 · 被引用 299 次
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 被引用 254 次
它引用的顶会 Paper19
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- MEGATRON-CNTRL: Controllable Story Generation with External Knowledge Using Large-Scale Language ModelsPeng Xu, Mostofa Patwary, Mohammad Shoeybi, Raul Puri 等EMNLP 2020 · 被引用 104 次
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong 等ACL 2020 · 被引用 103 次
- PlotMachines: Outline-Conditioned Generation with Dynamic Plot State TrackingHannah Rashkin, Asli Celikyilmaz, Yejin Choi, Jianfeng GaoEMNLP 2020 · 被引用 100 次
- With Little Power Comes Great ResponsibilityDallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia 等EMNLP 2020 · 被引用 76 次
相关 Paper
- Incorporating Worker Perspectives into MTurk Annotation Practices for NLPOlivia Huang, Eve Fleisig, Dan KleinEMNLP 2023 · 被引用 1 次
- A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for SummarizationLining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch 等ACL 2023 · 被引用 6 次
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu 等ACL 2021
- Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text GenerationKatelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi, Zongwan Cao 等ACL 2026
- GENIE: Toward Reproducible and Standardized Human Evaluation for Text GenerationDaniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie 等EMNLP 2022 · 被引用 13 次
