VerifyMatch: A Semi-Supervised Learning Paradigm for Natural Language Inference with Confidence-Aware MixUp
Seoyeon Park, Cornelia Caragea
Abstract
While natural language inference (NLI) has emerged as a prominent task for evaluating a model’s capability to perform natural language understanding, creating large benchmarks for training deep learning models imposes a significant challenge since it requires extensive human annotations. To overcome this, we propose to construct pseudo-generated samples (premise-hypothesis pairs) using class-specific fine-tuned large language models (LLMs) thereby reducing the human effort and the costs in annotating large amounts of data. However, despite the impressive performance of LLMs, it is necessary to verify that the pseudo-generated labels are actually correct. Towards this goal, in this paper, we propose VerifyMatch, a semi-supervised learning (SSL) approach in which the LLM pseudo-labels guide the training of the SSL model and, at the same time, the SSL model acts as a verifier of the LLM-generated data. In our approach, we retain all pseudo-labeled samples, but to ensure unlabeled data quality, we further propose to use MixUp whenever the verifier does not agree with the LLM-generated label or when they both agree on the label but the verifier has a low confidence—lower than an adaptive confidence threshold. We achieve competitive accuracy compared to strong baselines for NLI datasets in low-resource settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d06f590a-492f-4886-bf2a-a033a9aba7c9Cited by top-tier papers3
- MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text ClassificationIustin Sirbu, Robert-Adrian Popovici, Cornelia Caragea, Stefan Trausan-Matu et al.EMNLP 2025 · 1 citation
- CoMRes: Semi-Supervised Time Series Forecasting Utilizing Consensus Promotion of Multi-ResolutionYunju Cho, Jay-Yoon LeeICLR 2025
- LLM-Guided Co-Training for Text ClassificationMd Mezbaur Rahman, Cornelia CarageaEMNLP 2025
Builds on18
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
Related papers
- RegMixMatch: Optimizing Mixup Utilization in Semi-Supervised LearningHaorong Han, Jidong Yuan, Chixuan Wei, Zhongyang YuAAAI 2025 · 7 citations
- HyperMatch: Noise-Tolerant Semi-Supervised Learning via Relaxed Contrastive ConstraintBeitong Zhou, Jing Lu, Kerui Liu, Yunlu Xu et al.CVPR 2023
- BEACON: Budget-Aware Entity Matching Across DomainsNicholas Pulsone, Roee Shraga, Gregory GorenSIGMOD 2026 · 2 citations
- Debiased Self-Training for Semi-Supervised LearningBaixu Chen, Junguang Jiang, Ximei Wang, Pengfei Wan et al.NeurIPS 2022 · 162 citations
- Don't fear the unlabelled: safe semi-supervised learning via debiasingHugo Schmutz, Olivier Humbert, Pierre-Alexandre MatteiICLR 2023 · 1 citation
