AAAI2022

Scaling Up Influence Functions

Andrea Schioppa, Polina Zablotskaia, David Vilar, Artem Sokolov

被引用 149 次

摘要

We address efficient calculation of influence functions (Koh and Liang 2017) for tracking predictions back to the training data. We propose and analyze a new approach to speeding up the inverse Hessian calculation based on Arnoldi iteration (Arnoldi 1951) . With this improvement, we achieve, to the best of our knowledge, the first successful implementation of influence functions that scales to full-size (language and vision) Transformer models with several hundreds of millions of parameters. We evaluate our approach on image classification and sequence-to-sequence tasks with tens to a hundred of millions of training examples. Our code will be available at https://github.com/google-research/jax-influence .