Active Sampling for Text Classification with Subinstance Level Queries
Shayok Chakraborty, Ankita Singh
Abstract
Active learning algorithms are effective in identifying the salient and exemplar samples from large amounts of unlabeled data. This tremendously reduces the human annotation effort in inducing a machine learning model as only a few samples, which are identified by the algorithm, need to be labeled manually. In problem domains like text mining and video classification, human oracles peruse the data instances incrementally to derive an opinion about their class labels (such as reading a movie review progressively to assess its sentiment). In such applications, it is not necessary for the human oracles to review an unlabeled sample end-to-end in order to provide a label; it may be more efficient to identify an optimal subinstance size (percentage of the sample from the start) for each unlabeled sample, and request the human annotator to label the sample by analyzing only the subinstance, instead of the whole data sample. In this paper, we propose a novel framework to address this challenging problem, in an effort to further reduce the labeling burden on the human oracles and utilize the available labeling budget more efficiently. We pose the sample and subinstance size selection as a constrained optimization problem and derive a linear programming relaxation to select a batch of exemplar samples, together with the optimal subinstance size of each, which can potentially augment maximal information to the underlying classification model. Our extensive empirical studies on six challenging datasets from the text mining domain corroborate the practical usefulness of our framework over competing baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5dcafdfd-32ca-4fbf-b5ee-c0f1077fc0c0Builds on3
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 662 citations
- Asking the Right Questions to the Right Users: Active Learning with Imperfect OraclesShayok ChakrabortyAAAI 2020 · 23 citations
- State-Relabeling Adversarial Active LearningBeichen Zhang, Liang Li, Shijie Yang, Shuhui Wang et al.CVPR 2020
Related papers
- How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and BudgetGuy Hacohen, Daphna WeinshallNeurIPS 2023 · 23 citations
- Active Learning with Query Generation for Cost-Effective Text ClassificationYifan Yan, Sheng-Jun Huang, Shaoyi Chen, Meng Liao et al.AAAI 2020 · 27 citations
- On the Fragility of Active Learners for Text ClassificationAbhishek Ghose, Emma NguyenEMNLP 2024 · 2 citations
- Learning with Labeling Induced AbstentionsKareem Amin, Giulia DeSalvo, Afshin RostamizadehNeurIPS 2021 · 9 citations
- Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language ModelsChristopher Schröder, Gerhard HeyerEMNLP 2024 · 2 citations
