If in a Crowdsourced Data Annotation Pipeline, a GPT-4
Zeyu He, Chieh-Yang Huang, Chien-Kuang Cornelia Ding, Shaurya Rohatgi, Ting-Hao 'Kenneth' Huang
Abstract
Recent studies indicated GPT-4 outperforms online crowd workers in data labeling accuracy, notably workers from Amazon Mechanical Turk (MTurk). However, these studies were criticized for deviating from standard crowdsourcing practices and emphasizing individual workers' performances over the whole data-annotation process. This paper compared GPT-4 and an ethical and well-executed MTurk pipeline, with 415 workers labeling 3,177 sentence segments from 200 scholarly articles using the CODA-19 scheme. Two worker interfaces yielded 127,080 labels, which were then used to infer the final labels through eight label-aggregation algorithms. Our evaluation showed that despite best practices, MTurk pipeline's highest accuracy was 81.5%, whereas GPT-4 achieved 83.6%. Interestingly, when combining GPT-4's labels with crowd labels collected via an advanced worker interface for aggregation, 2 out of the 8 algorithms achieved an even higher accuracy (87.5%, 87.0%). Further analysis suggested that, when the crowd's and GPT-4's labeling strengths are complementary, aggregating them could increase labeling accuracy. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1116abb4-dd36-4145-b17d-29d27ac815f4Cited by top-tier papers13
- Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily AssistantGaole He, Gianluca Demartini, Ujwal GadirajuCHI 2025 · 91 citations
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas et al.CHI 2025 · 51 citations
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 33 citations
- Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are AbsentZeyu He, Saniya Naphade, Ting-Hao 'Kenneth' HuangCHI 2025 · 22 citations
- ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language ModelsYuzhe Gu, Ziwei Ji, Wenwei Zhang, Chengqi Lyu et al.NeurIPS 2024 · 20 citations
Builds on8
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
- Is GPT-3 a Good Data Annotator?Bosheng Ding, Chengwei Qin, Linlin Liu, Yew Ken Chia et al.ACL 2023 · 133 citations
- Adversarial Crowdsourcing Through Robust Rank-One Matrix CompletionQianqian Ma, Alex OlshevskyNeurIPS 2020 · 46 citations
- Attention Please: Your Attention Check Questions in Survey Studies Can Be Automatically AnsweredWeiping Pei, Arthur Mayer, Kaylynn Tu, Chuan YueWWW 2020 · 37 citations
Related papers
- A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for SummarizationLining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch et al.ACL 2023 · 6 citations
- Incorporating Worker Perspectives into MTurk Annotation Practices for NLPOlivia Huang, Eve Fleisig, Dan KleinEMNLP 2023 · 1 citation
- Reliance and Automation for Human-AI Collaborative Data Labeling Conflict ResolutionMichelle Brachman, Zahra Ashktorab, Michael Desmond, Evelyn Duesterwald et al.CSCW 2022 · 16 citations
- AI-Assisted Human Labeling: Batching for Efficiency without OverrelianceZahra Ashktorab, Michael Desmond, Josh Andres, Michael J. Muller et al.CSCW 2021 · 45 citations
- What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt et al.ACL 2021
