Co-training for Low Resource Scientific Natural Language Inference
Mobashir Sadat, Cornelia Caragea
摘要
Scientific Natural Language Inference (NLI) is the task of predicting the semantic relation between a pair of sentences extracted from research articles. The automatic annotation method based on distant supervision for the training set of SCINLI (Sadat and Caragea, 2022b), the first and most popular dataset for this task, results in label noise which inevitably degenerates the performance of classifiers. In this paper, we propose a novel co-training method that assigns weights based on the training dynamics of the classifiers to the distantly supervised labels, reflective of the manner they are used in the subsequent training epochs. That is, unlike the existing semi-supervised learning (SSL) approaches, we consider the historical behavior of the classifiers to evaluate the quality of the automatically annotated labels. Furthermore, by assigning importance weights instead of filtering out examples based on an arbitrary threshold on the predicted confidence, we maximize the usage of automatically labeled data, while ensuring that the noisy labels have a minimal impact on model training. The proposed method obtains an improvement of 1.5% in Macro F1 over the distant supervision baseline, and substantial improvements over several other strong SSL baselines. We make our code and data available on Github. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text ClassificationIustin Sirbu, Robert-Adrian Popovici, Cornelia Caragea, Stefan Trausan-Matu 等EMNLP 2025 · 被引用 1 次
- LLM-Guided Co-Training for Text ClassificationMd Mezbaur Rahman, Cornelia CarageaEMNLP 2025
它引用的顶会 Paper9
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Dash: Semi-Supervised Learning with Dynamic ThresholdingYi Xu, Lei Shang, Jinxing Ye, Qi Qian 等ICML 2021 · 被引用 287 次
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 被引用 182 次
- SciNLI: A Corpus for Natural Language Inference on Scientific TextMobashir Sadat, Cornelia CarageaACL 2022 · 被引用 41 次
- Simple Conversational Data Augmentation for Semi-supervised Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2021 · 被引用 31 次
相关 Paper
- SENT: Sentence-level Distant Relation Extraction via Negative TrainingRuotian Ma, Tao Gui, Linyang Li, Qi Zhang 等ACL 2021
- All Labels Are Not Created Equal: Enhancing Semi-Supervision via Label Grouping and Co-TrainingIslam Nassar, Samitha Herath, Ehsan Abbasnejad, Wray L. Buntine 等CVPR 2021
- Are Noisy Sentences Useless for Distant Supervised Relation Extraction?Yuming Shang, He Yan Huang, Xianling Mao, Xin Sun 等AAAI 2020 · 被引用 39 次
- Improving Distantly Supervised Relation Extraction by Natural Language InferenceKang Zhou, Qiao Qiao, Yuepei Li, Qi LiAAAI 2023 · 被引用 12 次
- Co-learning: Learning from Noisy Labels with Self-supervisionCheng Tan, Jun Xia, Lirong Wu, Stan Z. LiACM MM 2021 · 被引用 145 次
