Recovering Top-Two Answers and Confusion Probability in Multi-Choice Crowdsourcing
Hyeonsu Jeong, Hye Won Chung
摘要
Crowdsourcing has emerged as an effective platform for labeling large amounts of data in a cost- and time-efficient manner. Most previous work has focused on designing an efficient algorithm to recover only the ground-truth labels of the data. In this paper, we consider multi-choice crowdsourcing tasks with the goal of recovering not only the ground truth, but also the most confusing answer and the confusion probability. The most confusing answer provides useful information about the task by revealing the most plausible answer other than the ground truth and how plausible it is. To theoretically analyze such scenarios, we propose a model in which there are the top two plausible answers for each task, distinguished from the rest of the choices. Task difficulty is quantified by the probability of confusion between the top two, and worker reliability is quantified by the probability of giving an answer among the top two. Under this model, we propose a two-stage inference algorithm to infer both the top two answers and the confusion probability. We show that our algorithm achieves the minimax optimal convergence rate. We conduct both synthetic and real data experiments and demonstrate that our algorithm outperforms other recent algorithms. We also show the applicability of our algorithms in inferring the difficulty of tasks and in training neural networks with top-two soft labels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 被引用 362 次
- Adversarial Crowdsourcing Through Robust Rank-One Matrix CompletionQianqian Ma, Alex OlshevskyNeurIPS 2020 · 被引用 46 次
- Detecting and Preventing Confused Labels in Crowdsourced DataEvgeny Krivosheev, Siarhei Bykau, Fabio Casati, Sunil PrabhakarVLDB 2020 · 被引用 12 次
相关 Paper
- Semi-Supervised Multi-Label Learning from Crowds via Deep Sequential Generative ModelWanli Shi, Victor S. Sheng, Xiang Li, Bin GuKDD 2020 · 被引用 12 次
- Crowdsourcing System for Numerical Tasks based on Latent Topic Aware Worker ReliabilityZhuan Shi, Shanyang Jiang, Lan Zhang, Yang Du 等INFOCOM 2021 · 被引用 15 次
- Amortized Variational Inference for Partial-Label Learning: A Probabilistic Approach to Label DisambiguationTobias Fuchs, Nadja KleinICML 2026
- Origins of Algorithmic Instabilities in Crowdsourced RankingKeith Burghardt, Tad Hogg, Raissa M. D'Souza, Kristina Lerman 等CSCW 2020 · 被引用 4 次
- Eliciting Confidence for Improving Crowdsourced Audio AnnotationsAna Elisa Méndez Méndez, Mark Cartwright, Juan Pablo Bello, Oded NovCSCW 2022 · 被引用 10 次
