Debiasing Made State-of-the-art: Revisiting the Simple Seed-based Weak Supervision for Text Classification
Chengyu Dong, Zihan Wang, Jingbo Shang
Abstract
Recent advances in weakly supervised text classification mostly focus on designing sophisticated methods to turn high-level human heuristics into quality pseudo-labels. In this paper, we revisit the seed matching-based method, which is arguably the simplest way to generate pseudo-labels, and show that its power was greatly underestimated. We show that the limited performance of seed matching is largely due to the label bias injected by the simple seed-match rule, which prevents the classifier from learning reliable confidence for selecting high-quality pseudo-labels. Interestingly, simply deleting the seed words present in the matched input texts can mitigate the label bias and help learn better confidence. Subsequently, the performance achieved by seed matching can be improved significantly, making it on par with or even better than the state-of-theart. Furthermore, to handle the case when the seed words are not made known, we propose to simply delete the word tokens in the input text randomly with a high deletion ratio. Remarkably, seed matching equipped with this random deletion method can often achieve even better performance than that with seed deletion. We refer to our method as SimSeed, which is publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68d4d233-151d-40ae-8f4c-ec5cf800d065Cited by top-tier papers1
Ask how each one uses itBuilds on8
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- O2U-Net: A Simple Noisy Label Detection Approach for Deep Neural NetworksJinchi Huang, Lie Qu, Rongfei Jia, Binqiang ZhaoICCV 2019 · 276 citations
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong et al.EMNLP 2020 · 203 citations
Related papers
- Contextualized Weak Supervision for Text ClassificationDheeraj Mekala, Jingbo ShangACL 2020 · 121 citations
- META: Metadata-Empowered Weak Supervision for Text ClassificationDheeraj Mekala, Xinyang Zhang, Jingbo ShangEMNLP 2020 · 34 citations
- SoftMatch: Addressing the Quantity-Quality Tradeoff in Semi-supervised LearningHao Chen, Ran Tao, Yue Fan, Yidong Wang et al.ICLR 2023
- PIEClass: Weakly-Supervised Text Classification with Prompting and Noise-Robust Iterative Ensemble TrainingYunyi Zhang, Minhao Jiang, Yu Meng, Yu Zhang et al.EMNLP 2023 · 16 citations
- JointMatch: A Unified Approach for Diverse and Collaborative Pseudo-Labeling to Semi-Supervised Text ClassificationHenry Peng Zou, Cornelia CarageaEMNLP 2023 · 13 citations
