On Coresets for End-to-end Learning from Crowds
Hang Yang, Zhiwu Li, Witold Pedrycz
摘要
Crowdsourcing is a common approach for training datahungry models by collecting high-quality labeled data with human labor. With crowdsourcing data, the end-to-end learning paradigm is rising, where the classifier is concatenated with annotator-specific confusion layers and the two parts are co-trained in a parameter-coupled manner. However, learning with the size of a very large set of annotations is a challenge when computation or energy is limited. In this paper, we analyze and refine the coresets for end-to-end learning from crowds under the sensitivity sampling framework. This coreset is a small possible subset of annotations, so one can efficiently optimize the Coupled Cross-Entropy Minimization problem with guaranteed approximation. We first prove the lower bound, which shows no coresets smaller than complete data with confusion layers. Then, with workers' transition matrices Ar, we show that with the regularization term log det A ⊤ r Ar, this lower bound can be prevented. Our main result is that under mild assumptions, a smaller coreset exists for the regularized Coupled Cross-Entropy Minimization problem. An upper bound of sensitivity is proposed for designing a sampling algorithm called CrowdCore. The experimental results on synthetic and real-world datasets demonstrate the effectiveness of our analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Learning from Crowds by Modeling Common ConfusionsZhendong Chu, Jing Ma, Hongning WangAAAI 2021 · 被引用 60 次
- Generic Coreset for Scalable Learning of Monotonic Kernels: Logistic Regression, Sigmoid and moreElad Tolochinsky, Ibrahim Jubran, Dan FeldmanICML 2022 · 被引用 19 次
- Crowdsourcing via Annotator Co-occurrence Imputation and Provable Symmetric Nonnegative Matrix FactorizationShahana Ibrahim, Xiao FuICML 2021 · 被引用 12 次
- Coresets for Relational Data and The ApplicationsJiaxiang Chen, Qingyuan Yang, Ruomin Huang, Hu DingNeurIPS 2022 · 被引用 10 次
相关 Paper
- Deep Learning From Crowdsourced Labels: Coupled Cross-Entropy Minimization, Identifiability, and RegularizationShahana Ibrahim, Tri Nguyen, Xiao FuICLR 2023 · 被引用 3 次
- Coupled Confusion Correction: Learning from Crowds with Sparse AnnotationsHansong Zhang, Shikun Li, Dan Zeng, Chenggang Yan 等AAAI 2024 · 被引用 23 次
- No Dimensional Sampling Coresets for ClassificationMeysam Alishahi, Jeff M. PhillipsICML 2024 · 被引用 4 次
- Improve Learning from Crowds via Generative AugmentationZhendong Chu, Hongning WangKDD 2021 · 被引用 7 次
- Crowd Teaching with Imperfect LabelsYao Zhou, Arun Reddy Nelakurthi, Ross Maciejewski, Wei Fan 等WWW 2020 · 被引用 12 次
