Harnessing the Power of Beta Scoring in Deep Active Learning for Multi-Label Text Classification
Wei Tan, Ngoc Dang Nguyen, Lan Du, Wray L. Buntine
摘要
Within the scope of natural language processing, the domain of multi-label text classification is uniquely challenging due to its expansive and uneven label distribution. The complexity deepens due to the demand for an extensive set of annotated data for training an advanced deep learning model, especially in specialized fields where the labeling task can be labor-intensive and often requires domain-specific knowledge. Addressing these challenges, our study introduces a novel deep active learning strategy, capitalizing on the Beta family of proper scoring rules within the Expected Loss Reduction framework. It computes the expected increase in scores using the Beta Scoring Rules, which are then transformed into sample vector representations. These vector representations guide the diverse selection of informative sample, directly linking this process to the model's expected proper score. Comprehensive evaluations across both synthetic and real datasets reveal our method's capability to often outperform established acquisition techniques in multi-label text classification, presenting encouraging outcomes across various architectural and dataset scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford 等ICLR 2020 · 被引用 974 次
- Uncertainty-aware Active Learning for Optimal Bayesian ClassifierGuang Zhao, Edward R. Dougherty, Byung-Jun Yoon, Francis J. Alexander 等ICLR 2021 · 被引用 43 次
- Diversity Enhanced Active Learning with Strictly Proper Scoring RulesWei Tan, Lan Du, Wray L. BuntineNeurIPS 2021 · 被引用 40 次
- A Gaussian Process-Bayesian Bernoulli Mixture Model for Multi-Label Active LearningWeishi Shi, Dayou Yu, Qi YuNeurIPS 2021 · 被引用 10 次
- Imbalanced Semi-supervised Learning with Bias Adaptive ClassifierRenzhen Wang, Xixi Jia, Quanziang Wang, Yichen Wu 等ICLR 2023 · 被引用 4 次
相关 Paper
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch 等EMNLP 2020 · 被引用 144 次
- CoMAL: Contrastive Active Learning for Multi-Label Text ClassificationCheng Peng, Haobo Wang, Ke Chen, Lidan Shou 等KDD 2024 · 被引用 2 次
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 被引用 128 次
- Active Domain Adaptation via Clustering Uncertainty-weighted EmbeddingsViraj Prabhu, Arjun Chandrasekaran, Kate Saenko, Judy HoffmanICCV 2021 · 被引用 160 次
- Pretrained Generalized Autoregressive Model with Adaptive Probabilistic Label Clusters for Extreme Multi-label Text ClassificationHui Ye, Zhiyu Chen, Da-Han Wang, Brian D. DavisonICML 2020 · 被引用 57 次
