Learning Hyper Label Model for Programmatic Weak Supervision
Renzhi Wu, Shen-En Chen, Jieyu Zhang, Xu Chu
Abstract
To reduce the human annotation efforts, the programmatic weak supervision (PWS) paradigm abstracts weak supervision sources as labeling functions (LFs) and involves a label model to aggregate the output of multiple LFs to produce training labels. Most existing label models require a parameter learning step for each dataset. In this work, we present a hyper label model that (once learned) infers the ground-truth labels for each dataset in a single forward pass without dataset-specific parameter learning. The hyper label model approximates an optimal analytical (yet computationally intractable) solution of the ground-truth labels. We train the model on synthetic data generated in the way that ensures the model approximates the analytical optimal solution, and build the model upon Graph Neural Network (GNN) to ensure the model prediction being invariant (or equivariant) to the permutation of LFs (or data points). On 14 real-world datasets, our hyper label model outperforms the best existing methods in both accuracy (by 1.4 points on average) and efficiency (by six times on average). Our code is available at https://github.com/wurenzhi/hyper_label_model
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3eeab4c-8a80-438e-adc9-1b00092cf52bCited by top-tier papers7
- Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label ConfigurationsHao Chen, Ankit Shah, Jindong Wang, Ran Tao et al.NeurIPS 2024 · 22 citations
- Embroid: Unsupervised Prediction Smoothing Can Improve Few-Shot ClassificationNeel Guha, Mayee F. Chen, Kush Bhatia, Azalia Mirhoseini et al.NeurIPS 2023 · 6 citations
- Mitigating Source Bias for Fairer Weak SupervisionChangho Shin, Sonia Cromp, Dyah Adila, Frederic SalaNeurIPS 2023 · 5 citations
- Ground Truth Inference for Weakly Supervised Entity MatchingRenzhi Wu, Alexander Bendeck, Xu Chu, Yeye HeSIGMOD 2023 · 4 citations
- Statistical Analysis of an Adversarial Bayesian Weak Supervision MethodSteven AnNeurIPS 2025
Builds on8
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper et al.ICML 2020 · 130 citations
- Adversarial Multi Class Learning under Weak Supervision with Performance GuaranteesAlessio Mazzetto, Cyrus Cousins, Dylan Sam, Stephen H. Bach et al.ICML 2021 · 39 citations
- GOGGLES: Automatic Image Labeling with Affinity CodingNilaksh Das, Sanya Chaba, Renzhi Wu, Sakshi Gandhi et al.SIGMOD 2020 · 22 citations
- Learning to be a Statistician: Learned Estimator for Number of Distinct ValuesRenzhi Wu, Bolin Ding, Xu Chu, Zhewei Wei et al.VLDB 2022 · 16 citations
Related papers
- Label Propagation with Weak SupervisionRattana Pukdee, Dylan Sam, Pradeep Kumar Ravikumar, Nina BalcanICLR 2023
- A General Framework for Learning from Weak SupervisionHao Chen, Jindong Wang, Lei Feng, Xiang Li et al.ICML 2024 · 13 citations
- Refining Labeling Functions with Limited Labeled DataChenjie Li, Amir Gilad, Boris Glavic, Zhengjie Miao et al.KDD 2025 · 1 citation
- Understanding Programmatic Weak Supervision via Source-aware Influence FunctionJieyu Zhang, Haonan Wang, Cheng-Yu Hsieh, Alexander J. RatnerNeurIPS 2022 · 13 citations
- Characterizing the Impacts of Semi-supervised Learning for Weak SupervisionJeffrey Li, Jieyu Zhang, Ludwig Schmidt, Alexander J. RatnerNeurIPS 2023 · 9 citations
