Instance-dependent Early Stopping
Suqin Yuan, Runqi Lin, Lei Feng, Bo Han, Tongliang Liu
摘要
In machine learning practice, early stopping has been widely used to regularize models and can save computational costs by halting the training process when the model's performance on a validation set stops improving. However, conventional early stopping applies the same stopping criterion to all instances without considering their individual learning statuses, which leads to redundant computations on instances that are already well-learned. To further improve the efficiency, we propose an Instance-dependent Early Stopping (IES) method that adapts the early stopping mechanism from the entire training set to the instance level, based on the core principle that once the model has mastered an instance, the training on it should stop. IES considers an instance as mastered if the second-order differences of its loss value remain within a small range around zero. This offers a more consistent measure of an instance's learning status compared with directly using the loss value, and thus allows for a unified threshold to determine when an instance can be excluded from further backpropagation. We show that excluding mastered instances from backpropagation can increase the gradient norms, thereby accelerating the decrease of the training loss and speeding up the training process. Extensive experiments on benchmarks demonstrate that IES method can reduce backpropagation instances by 10%-50% while maintaining or even slightly improving the test accuracy and transfer learning performance of a model. Our implementation can be found at https://github.com/tmllab/2025_ICLR_IES .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Enhancing Sample Selection Against Label Noise by Cutting Mislabeled Easy ExamplesSuqin Yuan, Lei Feng, Bo Han, Tongliang LiuNeurIPS 2025 · 被引用 5 次
- Mitigating Mismatch within Reference-based Preference OptimizationSuqin Yuan, Xingrui Yu, Jiyang Zheng, Lei Feng 等ICLR 2026 · 被引用 4 次
- Progressive Data Dropout: An Embarrassingly Simple Approach to Train FasterShriram M. S, Xinyue Hao, Shihao Hou, Yang Lu 等NeurIPS 2025 · 被引用 2 次
- Samples Are Not Equal: A Sample Selection Approach for Deep ClusteringZhengxing Jiao, Yaxin Hou, Jun Ma, Yuhang Li 等ICLR 2026
它引用的顶会 Paper41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
相关 Paper
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 被引用 205 次
- Curriculum Learning by Dynamic Instance HardnessTianyi Zhou, Shengjie Wang, Jeff A. BilmesNeurIPS 2020 · 被引用 113 次
- MisDetect: Iterative Mislabel Detection using Early LossYuhao Deng, Chengliang Chai, Lei Cao, Nan Tang 等VLDB 2024 · 被引用 13 次
- Understanding and Improving Early Stopping for Learning with Noisy LabelsYingbin Bai, Erkun Yang, Bo Han, Yanhua Yang 等NeurIPS 2021 · 被引用 307 次
- Learning to Rank Learning CurvesMartin Wistuba, Tejaswini PedapatiICML 2020 · 被引用 31 次
