Understanding Influence Functions and Datamodels via Harmonic Analysis
Nikunj Saunshi, Arushi Gupta, Mark Braverman, Sanjeev Arora
摘要
Influence functions estimate effect of individual data points on predictions of the model on test data and were adapted to deep learning in Koh and Liang [2017] . They have been used for detecting data poisoning, detecting helpful and harmful examples, influence of groups of datapoints, etc. Recently, Ilyas et al. [2022] introduced a linear regression method they termed datamodels to predict the effect of training points on outputs on test data. The current paper seeks to provide a better theoretical understanding of such interesting empirical phenomena. The primary tool is harmonic analysis and the idea of noise stability. Contributions include: (a) Exact characterization of the learnt datamodel in terms of Fourier coefficients. (b) An efficient method to estimate the residual error and quality of the optimum linear datamodel without having to train the datamodel. (c) New insights into when influences of groups of datapoints may or may not add up linearly.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora 等ICML 2024 · 被引用 460 次
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc 等ICML 2023 · 被引用 260 次
- DsDm: Model-Aware Dataset Selection with DatamodelsLogan Engstrom, Axel Feldmann, Aleksander MadryICML 2024 · 被引用 105 次
- Most Influential Subset Selection: Challenges, Promises, and BeyondYuzheng Hu, Pingbang Hu, Han Zhao, Jiaqi W. MaNeurIPS 2024 · 被引用 39 次
- Stochastic Amortization: A Unified Approach to Accelerate Feature and Data AttributionIan Covert, Chanwoo Kim, Su-In Lee, James Y. Zou 等NeurIPS 2024 · 被引用 25 次
它引用的顶会 Paper14
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi 等NeurIPS 2022 · 被引用 185 次
- Scaling Up Influence FunctionsAndrea Schioppa, Polina Zablotskaia, David Vilar, Artem SokolovAAAI 2022 · 被引用 149 次
- Evaluation of Similarity-based ExplanationsKazuaki Hanawa, Sho Yokoi, Satoshi Hara, Kentaro InuiICLR 2021 · 被引用 79 次
相关 Paper
- Influence Functions in Deep Learning Are FragileSamyadeep Basu, Phillip Pope, Soheil FeiziICLR 2021 · 被引用 15 次
- RRInf: Efficient Influence Function Estimation via Ridge Regression for Large Language Models and Text-to-Image Diffusion ModelsZhuozhuo Tu, Cheng Chen, Yuxuan DuEMNLP 2025
- Characterizing the Influence of Graph ElementsZizhang Chen, Peizhao Li, Hongfu Liu, Pengyu HongICLR 2023 · 被引用 1 次
- The Mirrored Influence Hypothesis: Efficient Data Influence Estimation by Harnessing Forward PassesMyeongseob Ko, Feiyang Kang, Weiyan Shi, Ming Jin 等CVPR 2024
- HyDRA: Hypergradient Data Relevance Analysis for Interpreting Deep Neural NetworksYuanyuan Chen, Boyang Li, Han Yu, Pengcheng Wu 等AAAI 2021 · 被引用 50 次
