Crowd Teaching with Imperfect Labels
Yao Zhou, Arun Reddy Nelakurthi, Ross Maciejewski, Wei Fan, Jingrui He
摘要
The need for annotated labels to train machine learning models led to a surge in crowdsourcing - collecting labels from non-experts. Instead of annotating from scratch, given an imperfect labeled set, how can we leverage the label information obtained from amateur crowd workers to improve the data quality? Furthermore, is there a way to teach the amateur crowd workers using this imperfect labeled set in order to improve their labeling performance? In this paper, we aim to answer both questions via a novel interactive teaching framework, which uses visual explanations to simultaneously teach and gauge the confidence level of the crowd workers. Motivated by the huge demand for fine-grained label information in real-world applications, we start from the realistic and yet challenging assumption that neither the teacher nor the crowd workers are perfect. Then, we propose an adaptive scheme that could improve both of them through a sequence of interactions: the teacher teaches the workers using labeled data, and in return, the workers provide labels and the associated confidence level based on their own expertise. In particular, the teacher performs teaching using an empirical risk minimizer learned from an imperfect labeled set; the workers are assumed to have a forgetting behavior during learning and their learning rate depends on the interpretation difficulty of the teaching item. Furthermore, depending on the level of confidence when the workers perform labeling, we also show that the empirical risk minimizer used by the teacher is a reliable and realistic substitute of the unknown target concept by utilizing the unbiased surrogate loss. Finally, the performance of the proposed framework is demonstrated through experiments on multiple real-world image and text data sets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Local Clustering in Contextual Multi-Armed BanditsYikun Ban, Jingrui HeWWW 2021 · 被引用 51 次
- Iterative Teaching by Label SynthesisWeiyang Liu, Zhen Liu, Hanchen Wang, Liam Paull 等NeurIPS 2021 · 被引用 18 次
- Locality Sensitive TeachingZhaozhuo Xu, Beidi Chen, Chaojian Li, Weiyang Liu 等NeurIPS 2021 · 被引用 18 次
- Generic Outlier Detection in Multi-Armed BanditYikun Ban, Jingrui HeKDD 2020 · 被引用 17 次
- Nonparametric Iterative Machine TeachingChen Zhang, Xiaofeng Cao, Weiyang Liu, Ivor W. Tsang 等ICML 2023 · 被引用 13 次
相关 Paper
- Towards Professional Level Crowd Annotation of Expert Domain DataPei Wang, Nuno VasconcelosCVPR 2023
- A Machine Teaching Framework for Scalable RecognitionPei Wang, Nuno VasconcelosICCV 2021 · 被引用 11 次
- Hierarchical Crowdsourcing for Data Labeling with Heterogeneous CrowdHaodi Zhang, Wenxi Huang, Zhenhan Su, Junyang Chen 等ICDE 2023 · 被引用 4 次
- Teaching Active Human LearnersZizhe Wang, Hailong SunAAAI 2021 · 被引用 1 次
- Teaching Humans When to Defer to a Classifier via ExemplarsHussein Mozannar, Arvind Satyanarayan, David A. SontagAAAI 2022 · 被引用 49 次
