Learning from weak labelers as constraints
Vishwajeet Agrawal, Rattana Pukdee, Maria-Florina Balcan, Pradeep Kumar Ravikumar
Abstract
We study programmatic weak supervision, where, in contrast to labeled data, we have access to weak labelers, each of which either abstains or provides noisy labels corresponding to any input. Most previous approaches typically employ latent generative models that model the joint distribution of the weak labels and the latent "true" label. The caveats are that this relies on assumptions that may not always hold in practice, such as conditional independence assumptions over the joint distribution of the weak labelers and the latent true label, and more general implicit inductive biases in the latent generative models. In this work, we consider a more explicit form of side information that can be leveraged to denoise the weak labeler, namely the bounds on the average error of the weak labelers. We then propose a novel but natural weak supervision objective that minimizes a regularization functional subject to satisfying these bounds. This turns out to be a difficult constrained optimization problem due to discontinuous accuracy bound constraints. We provide a continuous optimization formulation for this objective through an alternating minimization algorithm that iteratively computes soft pseudo labels on the unlabeled data satisfying the constraints while being close to the model, and then updates the model on these labels until all the constraints are satisfied. We follow this with a theoretical analysis of this approach and provide insights into its denoising effects in training discriminative models given multiple weak labelers. Finally, we demonstrate the superior performance and robustness of our method on a popular weak supervision benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19f1ada8-5cd6-4ebb-b7d2-7c1dff20cb3aBuilds on10
- Part-dependent Label Noise: Towards Instance-dependent Label NoiseXiaobo Xia, Tongliang Liu, Bo Han, Nannan Wang et al.NeurIPS 2020 · 329 citations
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper et al.ICML 2020 · 130 citations
- Learning from Rules Generalizing Labeled ExemplarsAbhijeet Awasthi, Sabyasachi Ghosh, Rasna Goyal, Sunita SarawagiICLR 2020 · 93 citations
- End-to-End Weak SupervisionSalva Rühling Cachay, Benedikt Boecking, Artur DubrawskiNeurIPS 2021 · 48 citations
- Cryptographic Hardness of Learning Halfspaces with Massart NoiseIlias Diakonikolas, Daniel Kane, Pasin Manurangsi, Lisheng RenNeurIPS 2022 · 35 citations
Related papers
- Generative Modeling Helps Weak Supervision (and Vice Versa)Benedikt Boecking, Nicholas Carl Roberts, Willie Neiswanger, Stefano Ermon et al.ICLR 2023 · 1 citation
- Learning Hyper Label Model for Programmatic Weak SupervisionRenzhi Wu, Shen-En Chen, Jieyu Zhang, Xu ChuICLR 2023 · 2 citations
- Weak Supervision Performance Evaluation via Partial IdentificationFelipe Maia Polo, Subha Maity, Mikhail Yurochkin, Moulinath Banerjee et al.NeurIPS 2024 · 6 citations
- Creating Training Sets via Weak Indirect SupervisionJieyu Zhang, Bohan Wang, Xiangchen Song, Yujing Wang et al.ICLR 2022 · 17 citations
- Robust Weak Supervision with Variational Auto-EncodersFrancesco Tonolini, Nikolaos Aletras, Yunlong Jiao, Gabriella KazaiICML 2023 · 7 citations
