On the Accuracy of Newton Step and Influence Function Data Attributions
Ittai Rubinstein, Samuel Hopkins
摘要
Data attribution estimates how a trained model would change if a subset of training points were removed, and is a central primitive for tasks such as interpretability, data valuation, and machine unlearning. Despite its widespread use, our theoretical understanding of key data attribution methods -- Influence Functions (IF) and a single Newton Step (NS) -- remains limited: existing guarantees heavily rely on global strong convexity and yield bounds with pessimistic dependence on the parameter dimension and the number of removed samples . We give a new analysis of IF and NS for convex ERM that replaces global assumptions with local conditions: it suffices that the loss is strongly convex and sufficiently smooth only in a neighborhood of the first Newton step. As a concrete validation, we analyze logistic regression with Gaussian features and show that our bounds capture the correct scaling up to polylogarithmic factors, yielding matching upper and lower bounds and explaining observed regimes in which NS is markedly more accurate than IF, thereby resolving open questions raised by (Koh et al., 2019).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 被引用 516 次
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc 等ICML 2023 · 被引用 260 次
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi 等NeurIPS 2022 · 被引用 185 次
- What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsSang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao 等NeurIPS 2025 · 被引用 112 次
- Datamodels: Understanding Predictions with Data and Data with PredictionsAndrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc 等ICML 2022 · 被引用 66 次
相关 Paper
- Rescaled Influence Functions: Accurate Data Attribution in High DimensionIttai Rubinstein, Samuel B. HopkinsNeurIPS 2025 · 被引用 3 次
- Gaussian certified unlearning in high dimensions: A hypothesis testing approachAaradhya Pandey, Arnab Auddy, Haolin Zou, Arian Maleki 等ICLR 2026 · 被引用 5 次
- Machine Unlearning of Features and LabelsAlexander Warnecke, Lukas Pirch, Christian Wressnegger, Konrad RieckNDSS 2023
- On the Robustness of Removal-Based Feature AttributionsChris Lin, Ian Covert, Su-In LeeNeurIPS 2023 · 被引用 25 次
- A Versatile Influence Function for Data Attribution with Non-Decomposable LossJunwei Deng, Weijing Tang, Jiaqi W. MaICML 2025
