On Coresets for End-to-end Learning from Crowds
Hang Yang, Zhiwu Li, Witold Pedrycz
Abstract
Crowdsourcing is a common approach for training datahungry models by collecting high-quality labeled data with human labor. With crowdsourcing data, the end-to-end learning paradigm is rising, where the classifier is concatenated with annotator-specific confusion layers and the two parts are co-trained in a parameter-coupled manner. However, learning with the size of a very large set of annotations is a challenge when computation or energy is limited. In this paper, we analyze and refine the coresets for end-to-end learning from crowds under the sensitivity sampling framework. This coreset is a small possible subset of annotations, so one can efficiently optimize the Coupled Cross-Entropy Minimization problem with guaranteed approximation. We first prove the lower bound, which shows no coresets smaller than complete data with confusion layers. Then, with workers' transition matrices Ar, we show that with the regularization term log det A ⊤ r Ar, this lower bound can be prevented. Our main result is that under mild assumptions, a smaller coreset exists for the regularized Coupled Cross-Entropy Minimization problem. An upper bound of sensitivity is proposed for designing a sampling algorithm called CrowdCore. The experimental results on synthetic and real-world datasets demonstrate the effectiveness of our analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Learning from Crowds by Modeling Common ConfusionsZhendong Chu, Jing Ma, Hongning WangAAAI 2021 · 60 citations
- Generic Coreset for Scalable Learning of Monotonic Kernels: Logistic Regression, Sigmoid and moreElad Tolochinsky, Ibrahim Jubran, Dan FeldmanICML 2022 · 19 citations
- Crowdsourcing via Annotator Co-occurrence Imputation and Provable Symmetric Nonnegative Matrix FactorizationShahana Ibrahim, Xiao FuICML 2021 · 12 citations
- Coresets for Relational Data and The ApplicationsJiaxiang Chen, Qingyuan Yang, Ruomin Huang, Hu DingNeurIPS 2022 · 10 citations
Related papers
- Deep Learning From Crowdsourced Labels: Coupled Cross-Entropy Minimization, Identifiability, and RegularizationShahana Ibrahim, Tri Nguyen, Xiao FuICLR 2023 · 3 citations
- Coupled Confusion Correction: Learning from Crowds with Sparse AnnotationsHansong Zhang, Shikun Li, Dan Zeng, Chenggang Yan et al.AAAI 2024 · 23 citations
- No Dimensional Sampling Coresets for ClassificationMeysam Alishahi, Jeff M. PhillipsICML 2024 · 4 citations
- Improve Learning from Crowds via Generative AugmentationZhendong Chu, Hongning WangKDD 2021 · 7 citations
- Crowd Teaching with Imperfect LabelsYao Zhou, Arun Reddy Nelakurthi, Ross Maciejewski, Wei Fan et al.WWW 2020 · 12 citations
