Co-training for Low Resource Scientific Natural Language Inference
Mobashir Sadat, Cornelia Caragea
Abstract
Scientific Natural Language Inference (NLI) is the task of predicting the semantic relation between a pair of sentences extracted from research articles. The automatic annotation method based on distant supervision for the training set of SCINLI (Sadat and Caragea, 2022b), the first and most popular dataset for this task, results in label noise which inevitably degenerates the performance of classifiers. In this paper, we propose a novel co-training method that assigns weights based on the training dynamics of the classifiers to the distantly supervised labels, reflective of the manner they are used in the subsequent training epochs. That is, unlike the existing semi-supervised learning (SSL) approaches, we consider the historical behavior of the classifiers to evaluate the quality of the automatically annotated labels. Furthermore, by assigning importance weights instead of filtering out examples based on an arbitrary threshold on the predicted confidence, we maximize the usage of automatically labeled data, while ensuring that the noisy labels have a minimal impact on model training. The proposed method obtains an improvement of 1.5% in Macro F1 over the distant supervision baseline, and substantial improvements over several other strong SSL baselines. We make our code and data available on Github. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text ClassificationIustin Sirbu, Robert-Adrian Popovici, Cornelia Caragea, Stefan Trausan-Matu et al.EMNLP 2025 · 1 citation
- LLM-Guided Co-Training for Text ClassificationMd Mezbaur Rahman, Cornelia CarageaEMNLP 2025
Builds on9
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Dash: Semi-Supervised Learning with Dynamic ThresholdingYi Xu, Lei Shang, Jinxing Ye, Qi Qian et al.ICML 2021 · 287 citations
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 182 citations
- SciNLI: A Corpus for Natural Language Inference on Scientific TextMobashir Sadat, Cornelia CarageaACL 2022 · 41 citations
- Simple Conversational Data Augmentation for Semi-supervised Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2021 · 31 citations
Related papers
- SENT: Sentence-level Distant Relation Extraction via Negative TrainingRuotian Ma, Tao Gui, Linyang Li, Qi Zhang et al.ACL 2021
- All Labels Are Not Created Equal: Enhancing Semi-Supervision via Label Grouping and Co-TrainingIslam Nassar, Samitha Herath, Ehsan Abbasnejad, Wray L. Buntine et al.CVPR 2021
- Are Noisy Sentences Useless for Distant Supervised Relation Extraction?Yuming Shang, He Yan Huang, Xianling Mao, Xin Sun et al.AAAI 2020 · 39 citations
- Improving Distantly Supervised Relation Extraction by Natural Language InferenceKang Zhou, Qiao Qiao, Yuepei Li, Qi LiAAAI 2023 · 12 citations
- Co-learning: Learning from Noisy Labels with Self-supervisionCheng Tan, Jun Xia, Lirong Wu, Stan Z. LiACM MM 2021 · 145 citations
