The Memory-Perturbation Equation: Understanding Model's Sensitivity to Data
Peter Nickl, Lu Xu, Dharmesh Tailor, Thomas Möllenhoff, Mohammad Emtiyaz Khan
Abstract
Understanding model's sensitivity to its training data is crucial but can also be challenging and costly, especially during training. To simplify such issues, we present the Memory-Perturbation Equation (MPE) which relates model's sensitivity to perturbation in its training data. Derived using Bayesian principles, the MPE unifies existing sensitivity measures, generalizes them to a wide-variety of models and algorithms, and unravels useful properties regarding sensitivities. Our empirical results show that sensitivity estimates obtained during training can be used to faithfully predict generalization on unseen test data. The proposed equation is expected to be useful for future research on robust and adaptive learning. * Equal contribution. Part of this work was carried out when Dharmesh Tailor was at RIKEN AIP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc931597-8b61-4f84-b2f4-adb00a381f96Cited by top-tier papers9
- Variational Learning is Effective for Large Deep NetworksYuesong Shen, Nico Daheim, Bai Cong, Peter Nickl et al.ICML 2024 · 53 citations
- Training Data Attribution via Approximate UnrollingJuhan Bae, Wu Lin, Jonathan Lorraine, Roger B. GrosseNeurIPS 2024 · 41 citations
- Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts TransformersAlbus Li, Matthew WickerICML 2026 · 2 citations
- Set-Valued Sensitivity Analysis of Deep Neural NetworksXin Wang, Feilong Wang, Xuegang (Jeff) BanAAAI 2025 · 1 citation
- Joint Model and Data Sparsification via the Marginal LikelihoodAlexander Timans, Thomas Moellenhoff, Christian Andersson Naesseth, Mohammad Emtiyaz Khan et al.ICML 2026
Builds on14
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen et al.NeurIPS 2021 · 508 citations
Related papers
- Transformers Learn Low Sensitivity Functions: Investigations and ImplicationsBhavya Vasudeva, Deqing Fu, Tianyi Zhou, Elliott Kau et al.ICLR 2025
- Leveraging Unlabeled Data to Track MemorizationMahsa Forouzesh, Hanie Sedghi, Patrick ThiranICLR 2023
- TPV: Parameter Perturbations Through the Lens of Test Prediction VarianceDevansh ArpitICML 2026
- Interpreting Robust Optimization via Adversarial Influence FunctionsZhun Deng, Cynthia Dwork, Jialiang Wang, Linjun ZhangICML 2020 · 13 citations
- Training-Free Uncertainty Estimation for Dense Regression: Sensitivity as a SurrogateLu Mi, Hao Wang, Yonglong Tian, Hao He et al.AAAI 2022 · 36 citations
