Early Stopping Based on Unlabeled Samples in Text Classification
Hongseok Choi, Dongha Choi, Hyunju Lee
Abstract
Early stopping, which is widely used to prevent overfitting, is generally based on a separate validation set. However, in low resource settings, validation-based stopping can be risky because a small validation set may not be sufficiently representative, and the reduction in the number of samples by validation split may result in insufficient samples for training. In this study, we propose an early stopping method that uses unlabeled samples. The proposed method is based on confidence and class distribution similarities. To further improve the performance, we present a calibration method to better estimate the class distribution of the unlabeled samples. The proposed method is advantageous because it does not require a separate validation set and provides a better stopping point by using a large unlabeled set. Extensive experiments are conducted on five text classification datasets and several stop-methods are compared. Our results show that the proposed model even performs better than using an additional validation set as well as the existing stop-methods, in both balanced and imbalanced data settings. Our code is available at https://github. com/DMCB-GIST/BUS-stop .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- Boosting Few-Shot Text Classification via Distribution EstimationHan Liu, Feng Zhang, Xiaotong Zhang, Siyang Zhao et al.AAAI 2023 · 19 citations
- Calibrating Pseudo-Labeling with Class Distribution for Semi-supervised Text ClassificationWeiyi Yang, Richong Zhang, Junfan Chen, Jiawei ShengEMNLP 2025
- Instance-Selection-Inspired Undersampling Strategies for Bias Reduction in Small and Large Language Models for Binary Text ClassificationGuilherme Fonseca, Washington Cunha, Gabriel Prenassi, Marcos André Gonçalves et al.ACL 2025
- Free Lunch for Few-shot Learning: Distribution CalibrationShuo Yang, Lu Liu, Min XuICLR 2021 · 378 citations
- United We Stand: Using Epoch-Wise Agreement of Ensembles to Combat OverfitUri Stern, Daniel Shwartz, Daphna WeinshallAAAI 2024
