What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?
Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman
Abstract
Crowdsourcing is widely used to create data for common natural language understanding tasks. Despite the importance of these datasets for measuring and refining model understanding of language, there has been little focus on the crowdsourcing methods used for collecting the datasets. In this paper, we compare the efficacy of interventions that have been proposed in prior work as ways of improving data quality. We use multiple-choice question answering as a testbed and run a randomized trial by assigning crowdworkers to write questions under one of four different data collection protocols. We find that asking workers to write explanations for their examples is an ineffective stand-alone strategy for boosting NLU example difficulty. However, we find that training crowdworkers, and then using an iterative process of collecting data, sending feedback, and qualifying workers based on expert judgments is an effective means of collecting challenging data. But using crowdsourced, instead of expert judgments, to qualify workers and send feedback does not prove to be effective. We observe that the data from the iterative protocol with expert assessments is more challenging by several measures. Notably, the humanmodel gap on the unanimous agreement portion of this data is, on average, twice as large as the gap for the baseline protocol data. * Equal contribution. † Work done while at New York University. 10 We use pretrained models distributed with HuggingFace Transformers (Wolf et al., 2020) . batch size of 8, learning rate of 1.0 × 10 -5 , and finetune the models using the Adam optimizer for 4 epochs on the RACE dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd793808-144c-4b03-ae1f-5aca7a2346f2Cited by top-tier papers10
- Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI ChallengesQiaosi Wang, Michael Madaio, Shaun K. Kane, Shivani Kapania et al.CHI 2023 · 90 citations
- WebQA: Multihop and Multimodal QAYingshan Chang, Guihong Cao, Mridu Narang, Jianfeng Gao et al.CVPR 2022 · 58 citations
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani et al.EMNLP 2022 · 56 citations
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe et al.ACL 2024 · 30 citations
- CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about NegationAbhilasha Ravichander, Matt Gardner, Ana MarasovicEMNLP 2022 · 16 citations
Builds on7
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Getting Closer to AI Complete Question Answering: A Set of Prerequisite Real TasksAnna Rogers, Olga Kovaleva, Matthew Downey, Anna RumshiskyAAAI 2020 · 141 citations
- TORQUE: A Reading Comprehension Dataset of Temporal Ordering QuestionsQiang Ning, Hao Wu, Rujun Han, Nanyun Peng et al.EMNLP 2020 · 79 citations
Related papers
- New Protocols and Negative Results for Textual Entailment Data CollectionSamuel R. Bowman, Jennimaria Palomaki, Livio Baldini Soares, Emily PitlerEMNLP 2020 · 3 citations
- Crowd Teaching with Imperfect LabelsYao Zhou, Arun Reddy Nelakurthi, Ross Maciejewski, Wei Fan et al.WWW 2020 · 12 citations
- Hierarchical Crowdsourcing for Data Labeling with Heterogeneous CrowdHaodi Zhang, Wenxi Huang, Zhenhan Su, Junyang Chen et al.ICDE 2023 · 4 citations
- Can The Crowd Identify Misinformation Objectively?: The Effects of Judgment Scale and Assessor's BackgroundKevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina et al.SIGIR 2020 · 2 citations
- What Makes Reading Comprehension Questions Difficult?Saku Sugawara, Nikita Nangia, Alex Warstadt, Samuel R. BowmanACL 2022 · 15 citations
