Crowd Teaching with Imperfect Labels
Yao Zhou, Arun Reddy Nelakurthi, Ross Maciejewski, Wei Fan, Jingrui He
Abstract
The need for annotated labels to train machine learning models led to a surge in crowdsourcing - collecting labels from non-experts. Instead of annotating from scratch, given an imperfect labeled set, how can we leverage the label information obtained from amateur crowd workers to improve the data quality? Furthermore, is there a way to teach the amateur crowd workers using this imperfect labeled set in order to improve their labeling performance? In this paper, we aim to answer both questions via a novel interactive teaching framework, which uses visual explanations to simultaneously teach and gauge the confidence level of the crowd workers. Motivated by the huge demand for fine-grained label information in real-world applications, we start from the realistic and yet challenging assumption that neither the teacher nor the crowd workers are perfect. Then, we propose an adaptive scheme that could improve both of them through a sequence of interactions: the teacher teaches the workers using labeled data, and in return, the workers provide labels and the associated confidence level based on their own expertise. In particular, the teacher performs teaching using an empirical risk minimizer learned from an imperfect labeled set; the workers are assumed to have a forgetting behavior during learning and their learning rate depends on the interpretation difficulty of the teaching item. Furthermore, depending on the level of confidence when the workers perform labeling, we also show that the empirical risk minimizer used by the teacher is a reliable and realistic substitute of the unknown target concept by utilizing the unbiased surrogate loss. Finally, the performance of the proposed framework is demonstrated through experiments on multiple real-world image and text data sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b497e7fa-9040-406c-95e2-566f394cb4fdCited by top-tier papers7
- Local Clustering in Contextual Multi-Armed BanditsYikun Ban, Jingrui HeWWW 2021 · 51 citations
- Iterative Teaching by Label SynthesisWeiyang Liu, Zhen Liu, Hanchen Wang, Liam Paull et al.NeurIPS 2021 · 18 citations
- Locality Sensitive TeachingZhaozhuo Xu, Beidi Chen, Chaojian Li, Weiyang Liu et al.NeurIPS 2021 · 18 citations
- Generic Outlier Detection in Multi-Armed BanditYikun Ban, Jingrui HeKDD 2020 · 17 citations
- Nonparametric Iterative Machine TeachingChen Zhang, Xiaofeng Cao, Weiyang Liu, Ivor W. Tsang et al.ICML 2023 · 13 citations
Related papers
- Towards Professional Level Crowd Annotation of Expert Domain DataPei Wang, Nuno VasconcelosCVPR 2023
- A Machine Teaching Framework for Scalable RecognitionPei Wang, Nuno VasconcelosICCV 2021 · 11 citations
- Hierarchical Crowdsourcing for Data Labeling with Heterogeneous CrowdHaodi Zhang, Wenxi Huang, Zhenhan Su, Junyang Chen et al.ICDE 2023 · 4 citations
- Teaching Active Human LearnersZizhe Wang, Hailong SunAAAI 2021 · 1 citation
- Teaching Humans When to Defer to a Classifier via ExemplarsHussein Mozannar, Arvind Satyanarayan, David A. SontagAAAI 2022 · 49 citations
