Characterizing Datapoints via Second-Split Forgetting
Pratyush Maini, Saurabh Garg, Zachary C. Lipton, J. Zico Kolter
Abstract
Researchers investigating example hardness have increasingly focused on the dynamics by which neural networks learn and forget examples throughout training. Popular metrics derived from these dynamics include (i) the epoch at which examples are first correctly classified; (ii) the number of times their predictions flip during training; and (iii) whether their prediction flips if they are held out. However, these metrics do not distinguish among examples that are hard for distinct reasons, such as membership in a rare subpopulation, being mislabeled, or belonging to a complex subpopulation. In this paper, we propose - (SSFT), a complementary metric that tracks the epoch (if any) after which an original training example is forgotten as the network is fine-tuned on a randomly held out partition of the data. Across multiple benchmark datasets and modalities, we demonstrate that examples are forgotten quickly, and seemingly examples are forgotten comparatively slowly. By contrast, metrics only considering the first split learning dynamics struggle to differentiate the two. At large learning rates, SSFT tends to be robust across architectures, optimizers, and random seeds. From a practical standpoint, the SSFT can (i) help to identify mislabeled samples, the removal of which improves generalization; and (ii) provide insights about failure modes. Through theoretical analysis addressing overparameterized linear models, we provide insights into how the observed phenomena may arise. Code for reproducing our experiments can be found here: https://github.com/pratyushmaini/ssft
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 231a86ea-41de-4390-98fd-41fbb3eeb5caCited by top-tier papers25
- Can Neural Network Memorization Be Localized?Pratyush Maini, Michael Curtis Mozer, Hanie Sedghi, Zachary Chase Lipton et al.ICML 2023 · 82 citations
- Mechanistic Mode ConnectivityEkdeep Singh Lubana, Eric J. Bigelow, Robert P. Dick, David Scott Krueger et al.ICML 2023 · 57 citations
- T-MARS: Improving Visual Representations by Circumventing Text Feature LearningPratyush Maini, Sachin Goyal, Zachary Chase Lipton, J. Zico Kolter et al.ICLR 2024 · 43 citations
- Early Stopping Against Label Noise Without Validation DataSuqin Yuan, Lei Feng, Tongliang LiuICLR 2024 · 39 citations
- Beyond Confidence: Reliable Models Should Also Consider AtypicalityMert Yüksekgönül, Linjun Zhang, James Y. Zou, Carlos GuestrinNeurIPS 2023 · 32 citations
Builds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 798 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
Related papers
- Measuring Forgetting of Memorized Training ExamplesMatthew Jagielski, Om Thakkar, Florian Tramèr, Daphne Ippolito et al.ICLR 2023 · 15 citations
- Curriculum Learning by Dynamic Instance HardnessTianyi Zhou, Shengjie Wang, Jeff A. BilmesNeurIPS 2020 · 113 citations
- Does Continual Learning Equally Forget All Parameters?Haiyan Zhao, Tianyi Zhou, Guodong Long, Jing Jiang et al.ICML 2023 · 21 citations
- Predicting the Susceptibility of Examples to Catastrophic ForgettingGuy Hacohen, Tinne TuytelaarsICML 2025
- Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-TuningZekai Lin, Chao Xue, Di Liang, Xingsheng Han et al.ACL 2026 · 2 citations
