Detecting and Preventing Confused Labels in Crowdsourced Data
Evgeny Krivosheev, Siarhei Bykau, Fabio Casati, Sunil Prabhakar
Abstract
Crowdsourcing is a challenging activity for many reasons, from task design to workers' training, identification of low-quality annotators, and many more. A particularly subtle form of error is due to confusion of observations, that is, crowd workers (including diligent ones) that confuse items of a class i with items of a class j, either because they are similar or because the task description has failed to explain the differences. In this paper we show that confusion of observations can be a frequent occurrence in many tasks, and that such confusions cause a significant loss in accuracy. As a consequence, confusion detection is of primary importance for crowdsourced data labeling and classification. To address this problem we introduce an algorithm for confusion detection that leverages an inference procedure based on Markov Chain Monte Carlo (MCMC) sampling. We evaluate the algorithm via both synthetic datasets and crowdsourcing experiments and show that it has high accuracy in confusion detection (up to 99%). We experimentally show that quality is significantly improved without sacrificing efficiency. Finally, we show that detecting confusion is important as it can alert task designers early in the crowdsourcing process and lead designers to modify the task or add specific training and information to reduce the occurrence of workers' confusion. We show that even simple modifications, such as alerting workers of the risk of confusion, can improve performance significantly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66fa0adc-0951-40ac-bb6f-30b68bbf952fCited by top-tier papers2
- Quality of Sentiment Analysis Tools: The Reasons of InconsistencyWissam Maamar Kouadri, Mourad Ouziri, Salima Benbernou, Karima Echihabi et al.VLDB 2021 · 10 citations
- Recovering Top-Two Answers and Confusion Probability in Multi-Choice CrowdsourcingHyeonsu Jeong, Hye Won ChungICML 2023 · 2 citations
Related papers
- CrowdAct: Achieving High-Quality Crowdsourced Datasets in Mobile Activity RecognitionNattaya Mairittha, Tittaya Mairittha, Paula Lago, Sozo InoueUbiComp 2021 · 15 citations
- Noisy Label Learning with Instance-Dependent Outliers: Identifiability via Crowd WisdomTri Nguyen, Shahana Ibrahim, Xiao FuNeurIPS 2024 · 14 citations
- Quality Control in Crowdsourcing based on Fine-Grained Behavioral FeaturesWeiping Pei, Zhiju Yang, Monchu Chen, Chuan YueCSCW 2021 · 8 citations
- Semi-Supervised Multi-Label Learning from Crowds via Deep Sequential Generative ModelWanli Shi, Victor S. Sheng, Xiang Li, Bin GuKDD 2020 · 12 citations
- The Challenge of Variable Effort Crowdsourcing and How Visible Gold Can HelpDanula Hettiachchi, Mike Schaekermann, Tristan McKinney, Matthew LeaseCSCW 2021 · 21 citations
