Rescaled Influence Functions: Accurate Data Attribution in High Dimension
Ittai Rubinstein, Samuel B. Hopkins
Abstract
How does the training data affect a model's behavior? This is the question we seek to answer with data attribution. The leading practical approaches to data attribution are based on influence functions (IF). IFs utilize a first-order Taylor approximation to efficiently predict the effect of removing a set of samples from the training set without retraining the model, and are used in a wide variety of machine learning applications. However, especially in the high-dimensional regime (# params # samples), they are often imprecise and tend to underestimate the effect of sample removals, even for simple models such as logistic regression. We present rescaled influence functions (RIF), a new tool for data attribution which can be used as a drop-in replacement for influence functions, with little computational overhead but significant improvement in accuracy. We compare IF and RIF on a range of real-world datasets, showing that RIFs offer significantly better predictions in practice, and present a theoretical analysis explaining this improvement. Finally, we present a simple class of data poisoning attacks that would fool IF-based detections but would be detected by RIF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23b438e9-69ff-4c8a-a91a-3a0796b95d07Cited by top-tier papers2
- On the Accuracy of Newton Step and Influence Function Data AttributionsIttai Rubinstein, Samuel HopkinsICML 2026
- On the Fragility of Data Attribution When Learning Is DistributedXian Gao, Bo Hui, MIN-TE SUN, Wei-Shinn KuICML 2026
Builds on13
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 516 citations
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc et al.ICML 2023 · 260 citations
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi et al.NeurIPS 2022 · 185 citations
Related papers
- Taming Hyperparameter Sensitivity in Data Attribution: Practical Selection Without Costly RetrainingWeiyi Wang, Junwei Deng, Yuzheng Hu, Shiyuan Zhang et al.NeurIPS 2025 · 4 citations
- Final-Model-Only Data Attribution with a Unifying View of Gradient-Based MethodsDennis Wei, Inkit Padhi, Soumya Ghosh, Amit Dhurandhar et al.NeurIPS 2025 · 6 citations
- A Versatile Influence Function for Data Attribution with Non-Decomposable LossJunwei Deng, Weijing Tang, Jiaqi W. MaICML 2025
- Resolving Training Biases via Influence-based Data RelabelingShuming Kong, Yanyan Shen, Linpeng HuangICLR 2022 · 71 citations
- Distributional Training Data Attribution: What do Influence Functions Sample?Bruno Kacper Mlodozeniec, Isaac Reid, Sam Power, David Krueger et al.NeurIPS 2025
