Characterizing Datapoints via Second-Split Forgetting
Pratyush Maini, Saurabh Garg, Zachary C. Lipton, J. Zico Kolter
摘要
Researchers investigating example hardness have increasingly focused on the dynamics by which neural networks learn and forget examples throughout training. Popular metrics derived from these dynamics include (i) the epoch at which examples are first correctly classified; (ii) the number of times their predictions flip during training; and (iii) whether their prediction flips if they are held out. However, these metrics do not distinguish among examples that are hard for distinct reasons, such as membership in a rare subpopulation, being mislabeled, or belonging to a complex subpopulation. In this paper, we propose - (SSFT), a complementary metric that tracks the epoch (if any) after which an original training example is forgotten as the network is fine-tuned on a randomly held out partition of the data. Across multiple benchmark datasets and modalities, we demonstrate that examples are forgotten quickly, and seemingly examples are forgotten comparatively slowly. By contrast, metrics only considering the first split learning dynamics struggle to differentiate the two. At large learning rates, SSFT tends to be robust across architectures, optimizers, and random seeds. From a practical standpoint, the SSFT can (i) help to identify mislabeled samples, the removal of which improves generalization; and (ii) provide insights about failure modes. Through theoretical analysis addressing overparameterized linear models, we provide insights into how the observed phenomena may arise. Code for reproducing our experiments can be found here: https://github.com/pratyushmaini/ssft
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Can Neural Network Memorization Be Localized?Pratyush Maini, Michael Curtis Mozer, Hanie Sedghi, Zachary Chase Lipton 等ICML 2023 · 被引用 82 次
- Mechanistic Mode ConnectivityEkdeep Singh Lubana, Eric J. Bigelow, Robert P. Dick, David Scott Krueger 等ICML 2023 · 被引用 57 次
- T-MARS: Improving Visual Representations by Circumventing Text Feature LearningPratyush Maini, Sachin Goyal, Zachary Chase Lipton, J. Zico Kolter 等ICLR 2024 · 被引用 43 次
- Early Stopping Against Label Noise Without Validation DataSuqin Yuan, Lei Feng, Tongliang LiuICLR 2024 · 被引用 39 次
- Beyond Confidence: Reliable Models Should Also Consider AtypicalityMert Yüksekgönül, Linjun Zhang, James Y. Zou, Carlos GuestrinNeurIPS 2023 · 被引用 32 次
它引用的顶会 Paper18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 被引用 798 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
相关 Paper
- Measuring Forgetting of Memorized Training ExamplesMatthew Jagielski, Om Thakkar, Florian Tramèr, Daphne Ippolito 等ICLR 2023 · 被引用 15 次
- Curriculum Learning by Dynamic Instance HardnessTianyi Zhou, Shengjie Wang, Jeff A. BilmesNeurIPS 2020 · 被引用 113 次
- Does Continual Learning Equally Forget All Parameters?Haiyan Zhao, Tianyi Zhou, Guodong Long, Jing Jiang 等ICML 2023 · 被引用 21 次
- Predicting the Susceptibility of Examples to Catastrophic ForgettingGuy Hacohen, Tinne TuytelaarsICML 2025
- Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-TuningZekai Lin, Chao Xue, Di Liang, Xingsheng Han 等ACL 2026 · 被引用 2 次
