All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, Noah A. Smith
Abstract
Human evaluations are typically considered the gold standard in natural language generation, but as models' fluency improves, how well can evaluators detect and judge machinegenerated text? We run a study assessing nonexperts' ability to distinguish between humanand machine-authored text (GPT2 and GPT3) in three domains (stories, news articles, and recipes). We find that, without training, evaluators distinguished between GPT3-and humanauthored text at random chance level. We explore three approaches for quickly training evaluators to better identify GPT3-authored text (detailed instructions, annotated examples, and paired examples) and find that while evaluators' accuracy improved up to 55%, it did not significantly improve across the three domains. Given the inconsistent results across text domains and the often contradictory reasons evaluators gave for their judgments, we examine the role untrained human evaluations play in NLG evaluation and provide recommendations to NLG researchers for improving human evaluations of text generated from state-of-the-art models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eec2ff6b-6571-4ac1-b5ec-6d833181d41bCited by top-tier papers87
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- QUARK: Controllable Text Generation with Reinforced UnlearningXiming Lu, Sean Welleck, Jack Hessel, Liwei Jiang et al.NeurIPS 2022 · 290 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer ReviewsWeixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp et al.ICML 2024 · 213 citations
- Working With AI to Persuade: Examining a Large Language Model's Ability to Generate Pro-Vaccination MessagesElise Karinshak, Sunny Xun Liu, Joon Sung Park, Jeffrey T. HancockCSCW 2023 · 163 citations
Builds on5
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Authorship Attribution for Neural Text GenerationAdaku Uchendu, Thai Le, Kai Shu, Dongwon LeeEMNLP 2020 · 110 citations
- With Little Power Comes Great ResponsibilityDallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia et al.EMNLP 2020 · 76 citations
- Automatic Detection of Generated Text is Easiest when Humans are FooledDaphne Ippolito, Daniel Duckworth, Chris Callison-Burch, Douglas EckACL 2020 · 21 citations
- STORIUM: A Dataset and Evaluation Platform for Machine-in-the-Loop Story GenerationNader Akoury, Shufan Wang, Josh Whiting, Stephen Hood et al.EMNLP 2020 · 5 citations
Related papers
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textJenna Russell, Marzena Karpinska, Mohit IyyerACL 2025 · 39 citations
- Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AIYuxia Wang, Rui Xing, Jonibek Mansurov, Giovanni Puccetti et al.ACL 2026 · 3 citations
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextXianjun Yang, Wei Cheng, Yue Wu, Linda Ruth Petzold et al.ICLR 2024 · 173 citations
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation ModelsWangchunshu Zhou, Ke XuAAAI 2020 · 49 citations
- Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated TextLiam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi et al.AAAI 2023 · 112 citations
