What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?
Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman
摘要
Crowdsourcing is widely used to create data for common natural language understanding tasks. Despite the importance of these datasets for measuring and refining model understanding of language, there has been little focus on the crowdsourcing methods used for collecting the datasets. In this paper, we compare the efficacy of interventions that have been proposed in prior work as ways of improving data quality. We use multiple-choice question answering as a testbed and run a randomized trial by assigning crowdworkers to write questions under one of four different data collection protocols. We find that asking workers to write explanations for their examples is an ineffective stand-alone strategy for boosting NLU example difficulty. However, we find that training crowdworkers, and then using an iterative process of collecting data, sending feedback, and qualifying workers based on expert judgments is an effective means of collecting challenging data. But using crowdsourced, instead of expert judgments, to qualify workers and send feedback does not prove to be effective. We observe that the data from the iterative protocol with expert assessments is more challenging by several measures. Notably, the humanmodel gap on the unanimous agreement portion of this data is, on average, twice as large as the gap for the baseline protocol data. * Equal contribution. † Work done while at New York University. 10 We use pretrained models distributed with HuggingFace Transformers (Wolf et al., 2020) . batch size of 8, learning rate of 1.0 × 10 -5 , and finetune the models using the Adam optimizer for 4 epochs on the RACE dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI ChallengesQiaosi Wang, Michael Madaio, Shaun K. Kane, Shivani Kapania 等CHI 2023 · 被引用 90 次
- WebQA: Multihop and Multimodal QAYingshan Chang, Guihong Cao, Mridu Narang, Jianfeng Gao 等CVPR 2022 · 被引用 58 次
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani 等EMNLP 2022 · 被引用 56 次
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe 等ACL 2024 · 被引用 30 次
- CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about NegationAbhilasha Ravichander, Matt Gardner, Ana MarasovicEMNLP 2022 · 被引用 16 次
它引用的顶会 Paper7
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Getting Closer to AI Complete Question Answering: A Set of Prerequisite Real TasksAnna Rogers, Olga Kovaleva, Matthew Downey, Anna RumshiskyAAAI 2020 · 被引用 141 次
- TORQUE: A Reading Comprehension Dataset of Temporal Ordering QuestionsQiang Ning, Hao Wu, Rujun Han, Nanyun Peng 等EMNLP 2020 · 被引用 79 次
相关 Paper
- New Protocols and Negative Results for Textual Entailment Data CollectionSamuel R. Bowman, Jennimaria Palomaki, Livio Baldini Soares, Emily PitlerEMNLP 2020 · 被引用 3 次
- Crowd Teaching with Imperfect LabelsYao Zhou, Arun Reddy Nelakurthi, Ross Maciejewski, Wei Fan 等WWW 2020 · 被引用 12 次
- Hierarchical Crowdsourcing for Data Labeling with Heterogeneous CrowdHaodi Zhang, Wenxi Huang, Zhenhan Su, Junyang Chen 等ICDE 2023 · 被引用 4 次
- Can The Crowd Identify Misinformation Objectively?: The Effects of Judgment Scale and Assessor's BackgroundKevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina 等SIGIR 2020 · 被引用 2 次
- What Makes Reading Comprehension Questions Difficult?Saku Sugawara, Nikita Nangia, Alex Warstadt, Samuel R. BowmanACL 2022 · 被引用 15 次
