Memorization Through the Lens of Curvature of Loss Function Around Samples
Isha Garg, Deepak Ravikumar, Kaushik Roy
Abstract
Deep neural networks are over-parameterized and easily overfit the datasets they train on. In the extreme case, it has been shown that these networks can memorize a training set with fully randomized labels. We propose using the curvature of loss function around each training sample, averaged over training epochs, as a measure of memorization of the sample. We use this metric to study the generalization versus memorization properties of different samples in popular image datasets and show that it captures memorization statistics well, both qualitatively and quantitatively. We first show that the high curvature samples visually correspond to long-tailed, mislabeled, or conflicting samples, those that are most likely to be memorized. This analysis helps us find, to the best of our knowledge, a novel failure mode on the CIFAR100 and ImageNet datasets: that of duplicated images with differing labels. Quantitatively, we corroborate the validity of our scores via two methods. First, we validate our scores against an independent and comprehensively calculated baseline, by showing high cosine similarity with the memorization scores released by Feldman and Zhang (2020). Second, we inject corrupted samples which are memorized by the network, and show that these are learned with high curvature. To this end, we synthetically mislabel a random subset of the dataset. We overfit a network to it and show that sorting by curvature yields high AUROC values for identifying the corrupted samples. An added advantage of our method is that it is scalable, as it requires training only a single network as opposed to the thousands trained by the baseline, while capturing the aforementioned failure mode that the baseline fails to identify.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 503911e5-29be-43b8-a641-ef37ba110fa7Cited by top-tier papers14
- NoiseGPT: Label Noise Detection and Rectification through Probability CurvatureHaoyu Wang, Zhuo Huang, Zhiwei Lin, Tongliang LiuNeurIPS 2024 · 27 citations
- Unveiling Privacy, Memorization, and Input Curvature LinksDeepak Ravikumar, Efstathia Soufleri, Abolfazl Hashemi, Kaushik RoyICML 2024 · 16 citations
- Remaining-data-free Machine Unlearning by Suppressing Sample ContributionXinwen Cheng, Zhehao Huang, Wenxing Zhou, Zhengbao He et al.ICLR 2026 · 11 citations
- SAP: Corrective Machine Unlearning with Scaled Activation Projection for Label Noise RobustnessSangamesh Kodge, Deepak Ravikumar, Gobinda Saha, Kaushik RoyAAAI 2025 · 10 citations
- Curvature Clues: Decoding Deep Learning Privacy with Input Loss CurvatureDeepak Ravikumar, Efstathia Soufleri, Kaushik RoyNeurIPS 2024 · 10 citations
Builds on10
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
Related papers
- Random Label Prediction Heads for Studying Memorization in Deep Neural NetworksMarlon Becker, Jonas Konrad, Luis Garcia Rodriguez, Benjamin RisseICLR 2026
- Towards Memorization Estimation: Fast, Formal and FreeDeepak Ravikumar, Efstathia Soufleri, Abolfazl Hashemi, Kaushik RoyICML 2025
- Memorization Through the Lens of Sample GradientsDeepak Ravikumar, Efstathia Soufleri, Abolfazl Hashemi, Kaushik RoyICLR 2026
- On Memorization in Probabilistic Deep Generative ModelsGerrit J. J. van den Burg, Christopher K. I. WilliamsNeurIPS 2021 · 92 citations
- Localizing Memorized Regions in Diffusion Models via Coordinate-Wise Curvature DifferencesGwangho Kim, Sungyoon LeeICML 2026
