If Influence Functions are the Answer, Then What is the Question?
Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi, Roger B. Grosse
Abstract
Influence functions efficiently estimate the effect of removing a single training data point on a model's learned parameters. While influence estimates align well with leave-one-out retraining for linear models, recent works have shown this alignment is often poor in neural networks. In this work, we investigate the specific factors that cause this discrepancy by decomposing it into five separate terms. We study the contributions of each term on a variety of architectures and datasets and how they vary with factors such as network width and training time. While practical influence function estimates may be a poor match to leave-one-out retraining for nonlinear networks, we show they are often a good approximation to a different object we term the proximal Bregman response function (PBRF). Since the PBRF can still be used to answer many of the questions motivating influence functions, such as identifying influential or mislabeled examples, our results suggest that current algorithms for influence function estimation give more informative results than previous error analyses would suggest.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1c505bc-f35b-4e4e-b0ea-5c5d497c864bCited by top-tier papers64
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc et al.ICML 2023 · 260 citations
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion ModelsYongchan Kwon, Eric Wu, Kevin Wu, James ZouICLR 2024 · 112 citations
- What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsSang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao et al.NeurIPS 2025 · 112 citations
- Dataset Distillation with Convexified Implicit GradientsNoel Loo, Ramin M. Hasani, Mathias Lechner, Daniela RusICML 2023 · 56 citations
Builds on13
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- On Warm-Starting Neural Network TrainingJordan T. Ash, Ryan P. AdamsNeurIPS 2020 · 288 citations
- Scaling Up Influence FunctionsAndrea Schioppa, Polina Zablotskaia, David Vilar, Artem SokolovAAAI 2022 · 149 citations
- Do GANs always have Nash equilibria?Farzan Farnia, Asuman E. OzdaglarICML 2020 · 93 citations
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence FunctionsXiaochuang Han, Byron C. Wallace, Yulia TsvetkovACL 2020 · 91 citations
Related papers
- Influence Functions for Edge Edits in Non-Convex Graph Neural NetworksJaeseung Heo, Kyeongheung Yun, Seokwon Yoon, MoonJeong Park et al.NeurIPS 2025 · 2 citations
- Theoretical and Practical Perspectives on what Influence Functions DoAndrea Schioppa, Katja Filippova, Ivan Titov, Polina ZablotskaiaNeurIPS 2023 · 38 citations
- Influence Functions in Deep Learning Are FragileSamyadeep Basu, Phillip Pope, Soheil FeiziICLR 2021 · 15 citations
- Towards Robust Influence Functions with Flat Validation MinimaXichen Ye, Yifan Wu, Weizhong Zhang, Cheng Jin et al.ICML 2025
- Efficient Parametric Approximations of Neural Network Function Space DistanceNikita Dhawan, Sicong Huang, Juhan Bae, Roger Baker GrosseICML 2023 · 7 citations
