Learning from weak labelers as constraints
Vishwajeet Agrawal, Rattana Pukdee, Maria-Florina Balcan, Pradeep Kumar Ravikumar
摘要
We study programmatic weak supervision, where, in contrast to labeled data, we have access to weak labelers, each of which either abstains or provides noisy labels corresponding to any input. Most previous approaches typically employ latent generative models that model the joint distribution of the weak labels and the latent "true" label. The caveats are that this relies on assumptions that may not always hold in practice, such as conditional independence assumptions over the joint distribution of the weak labelers and the latent true label, and more general implicit inductive biases in the latent generative models. In this work, we consider a more explicit form of side information that can be leveraged to denoise the weak labeler, namely the bounds on the average error of the weak labelers. We then propose a novel but natural weak supervision objective that minimizes a regularization functional subject to satisfying these bounds. This turns out to be a difficult constrained optimization problem due to discontinuous accuracy bound constraints. We provide a continuous optimization formulation for this objective through an alternating minimization algorithm that iteratively computes soft pseudo labels on the unlabeled data satisfying the constraints while being close to the model, and then updates the model on these labels until all the constraints are satisfied. We follow this with a theoretical analysis of this approach and provide insights into its denoising effects in training discriminative models given multiple weak labelers. Finally, we demonstrate the superior performance and robustness of our method on a popular weak supervision benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Part-dependent Label Noise: Towards Instance-dependent Label NoiseXiaobo Xia, Tongliang Liu, Bo Han, Nannan Wang 等NeurIPS 2020 · 被引用 329 次
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper 等ICML 2020 · 被引用 130 次
- Learning from Rules Generalizing Labeled ExemplarsAbhijeet Awasthi, Sabyasachi Ghosh, Rasna Goyal, Sunita SarawagiICLR 2020 · 被引用 93 次
- End-to-End Weak SupervisionSalva Rühling Cachay, Benedikt Boecking, Artur DubrawskiNeurIPS 2021 · 被引用 48 次
- Cryptographic Hardness of Learning Halfspaces with Massart NoiseIlias Diakonikolas, Daniel Kane, Pasin Manurangsi, Lisheng RenNeurIPS 2022 · 被引用 35 次
相关 Paper
- Generative Modeling Helps Weak Supervision (and Vice Versa)Benedikt Boecking, Nicholas Carl Roberts, Willie Neiswanger, Stefano Ermon 等ICLR 2023 · 被引用 1 次
- Learning Hyper Label Model for Programmatic Weak SupervisionRenzhi Wu, Shen-En Chen, Jieyu Zhang, Xu ChuICLR 2023 · 被引用 2 次
- Weak Supervision Performance Evaluation via Partial IdentificationFelipe Maia Polo, Subha Maity, Mikhail Yurochkin, Moulinath Banerjee 等NeurIPS 2024 · 被引用 6 次
- Creating Training Sets via Weak Indirect SupervisionJieyu Zhang, Bohan Wang, Xiangchen Song, Yujing Wang 等ICLR 2022 · 被引用 17 次
- Robust Weak Supervision with Variational Auto-EncodersFrancesco Tonolini, Nikolaos Aletras, Yunlong Jiao, Gabriella KazaiICML 2023 · 被引用 7 次
