Refining Labeling Functions with Limited Labeled Data
Chenjie Li, Amir Gilad, Boris Glavic, Zhengjie Miao, Sudeepa Roy
摘要
Programmatic weak supervision (PWS) significantly reduces human effort for labeling data by combining the outputs of userprovided labeling functions (LFs) on unlabeled datapoints. However, the quality of the generated labels depends directly on the accuracy of the LFs. In this work, we study the problem of fixing LFs based on a small set of labeled examples. Towards this goal, we develop novel techniques for repairing a set of LFs by minimally changing their results on the labeled examples such that the fixed LFs ensure that (i) there is sufficient evidence for the correct label of each labeled datapoint and (ii) the accuracy of each repaired LF is sufficiently high. We model LFs as conditional rules, which enables us to refine them, i.e., to selectively change their output for some inputs. We demonstrate experimentally that our system improves the quality of LFs based on surprisingly small sets of labeled datapoints. CCS Concepts • Information systems → Data cleaning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan 等VLDB 2024 · 被引用 165 次
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper 等ICML 2020 · 被引用 130 次
- Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data ProgrammingCheng-Yu Hsieh, Jieyu Zhang, Alexander J. RatnerVLDB 2022 · 被引用 17 次
- Adaptive Rule Discovery for Labeling Text DataSainyam Galhotra, Behzad Golshan, Wang-Chiew TanSIGMOD 2021 · 被引用 14 次
- Understanding Programmatic Weak Supervision via Source-aware Influence FunctionJieyu Zhang, Haonan Wang, Cheng-Yu Hsieh, Alexander J. RatnerNeurIPS 2022 · 被引用 13 次
相关 Paper
- Learning Hyper Label Model for Programmatic Weak SupervisionRenzhi Wu, Shen-En Chen, Jieyu Zhang, Xu ChuICLR 2023 · 被引用 2 次
- Characterizing the Impacts of Semi-supervised Learning for Weak SupervisionJeffrey Li, Jieyu Zhang, Ludwig Schmidt, Alexander J. RatnerNeurIPS 2023 · 被引用 9 次
- Robust Weak Supervision with Variational Auto-EncodersFrancesco Tonolini, Nikolaos Aletras, Yunlong Jiao, Gabriella KazaiICML 2023 · 被引用 7 次
- Statistical Analysis of an Adversarial Bayesian Weak Supervision MethodSteven AnNeurIPS 2025
- Learning from weak labelers as constraintsVishwajeet Agrawal, Rattana Pukdee, Maria-Florina Balcan, Pradeep Kumar RavikumarICLR 2025
