Adaptive Rule Discovery for Labeling Text Data
Sainyam Galhotra, Behzad Golshan, Wang-Chiew Tan
摘要
Creating and collecting labeled data is one of the major bottlenecks in machine learning pipelines and the emergence of automated feature generation techniques such as deep learning, which typically requires a lot of training data, has further exacerbated the problem. While weak-supervision techniques have circumvented this bottleneck, existing frameworks either require users to write a set of diverse, highquality rules to label data (e.g., Snorkel), or require a labeled subset of the data to automatically mine rules (e.g., Snuba). The process of manually writing rules can be tedious and time consuming. At the same time, creating a labeled subset of the data can be costly and even infeasible in imbalanced settings. This is due to the fact that a random sample in imbalanced settings often contains only a few positive instances.
To address these shortcomings, we present Darwin, an interactive system designed to alleviate the task of writing rules for labeling text data in weakly-supervised settings. Given an initial labeling rule, Darwin automatically generates a set of candidate rules for the labeling task at hand, and utilizes the annotator's feedback to adapt the candidate rules. We describe how Darwin is scalable and versatile. It can operate over large text corpora (i.e., more than 1 million sentences) and supports a wide range of labeling functions (i.e., any function that can be specified using a context free grammar). Finally, we demonstrate with a suite of experiments over five real-world datasets that Darwin enables annotators to generate weakly-supervised labels efficiently and with a small cost. In fact, our experiments show that rules discovered by Darwin on average identify 40% more positive instances compared to Snuba even when it is provided with 1000 labeled instances.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Prompt-Based Rule Discovery and Boosting for Interactive Weakly-Supervised LearningRongzhi Zhang, Yue Yu, Pranav Shetty, Le Song 等ACL 2022 · 被引用 29 次
- Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data ProgrammingCheng-Yu Hsieh, Jieyu Zhang, Alexander J. RatnerVLDB 2022 · 被引用 17 次
- Self-Supervised Self-Supervision by Combining Deep Learning and Probabilistic LogicHunter Lang, Hoifung PoonAAAI 2021 · 被引用 13 次
- Witan: Unsupervised Labelling Function Generation for Assisted Data ProgrammingBenjamin Denham, Edmund M.-K. Lai, Roopak Sinha, Muhammad Asif NaeemVLDB 2022 · 被引用 12 次
- Refining Labeling Functions with Limited Labeled DataChenjie Li, Amir Gilad, Boris Glavic, Zhengjie Miao 等KDD 2025 · 被引用 1 次
相关 Paper
- Interactive Weak Supervision: Learning Useful Heuristics for Data LabelingBenedikt Boecking, Willie Neiswanger, Eric P. Xing, Artur DubrawskiICLR 2021 · 被引用 8 次
- Ground Truth Inference for Weakly Supervised Entity MatchingRenzhi Wu, Alexander Bendeck, Xu Chu, Yeye HeSIGMOD 2023 · 被引用 4 次
- DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language ModelsGuozheng Li, Ao Wang, Shaoxiang Wang, Yu Zhang 等CHI 2026
- Inspector Gadget: A Data Programming-based Labeling System for Industrial ImagesGeon Heo, Yuji Roh, Seonghyeon Hwang, Dayun Lee 等VLDB 2021 · 被引用 9 次
- Robust Weak Supervision with Variational Auto-EncodersFrancesco Tonolini, Nikolaos Aletras, Yunlong Jiao, Gabriella KazaiICML 2023 · 被引用 7 次
