Explaining Neural Matrix Factorization with Gradient Rollback
Carolin Lawrence, Timo Sztyler, Mathias Niepert
Abstract
Explaining the predictions of neural black-box models is an important problem, especially when such models are used in applications where user trust is crucial. Estimating the influence of training examples on a learned neural model's behavior allows us to identify training examples most responsible for a given prediction and, therefore, to faithfully explain the output of a black-box model. The most generally applicable existing method is based on influence functions, which scale poorly for larger sample sizes and models.
We propose gradient rollback, a general approach for influence estimation, applicable to neural models where each parameter update step during gradient descent touches a smaller number of parameters, even if the overall number of parameters is large. Neural matrix factorization models trained with gradient descent are part of this model class. These models are popular and have found a wide range of applications in industry. Especially knowledge graph embedding methods, which belong to this class, are used extensively. We show that gradient rollback is highly efficient at both training and test time. Moreover, we show theoretically that the difference between gradient rollback's influence approximation and the true influence on a model's behavior is smaller than known bounds on the stability of stochastic gradient descent. This establishes that gradient rollback is robustly estimating example influence. We also conduct experiments which show that gradient rollback provides faithful explanations for knowledge base completion and recommender datasets. An implementation and an appendix are available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bb3e457-5f43-4707-8c83-53aa82a6ac8dCited by top-tier papers3
- Adversarial Attacks on Knowledge Graph Embeddings via Instance Attribution MethodsPeru Bhardwaj, John D. Kelleher, Luca Costabello, Declan O'SullivanEMNLP 2021 · 17 citations
- Generating and Evaluating Plausible Explanations for Knowledge Graph CompletionAntonio Di Mauro, Zhao Xu, Wiem Ben Rim, Timo Sztyler et al.ACL 2024 · 2 citations
- Poisoning Knowledge Graph Embeddings via Relation Inference PatternsPeru Bhardwaj, John D. Kelleher, Luca Costabello, Declan O'SullivanACL 2021
Builds on1
Related papers
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Right for Better Reasons: Training Differentiable Models by Constraining their Influence FunctionsXiaoting Shao, Arseny Skryagin, Wolfgang Stammer, Patrick Schramowski et al.AAAI 2021 · 45 citations
- On Second-Order Group Influence Functions for Black-Box PredictionsSamyadeep Basu, Xuchen You, Soheil FeiziICML 2020 · 28 citations
- Bayesian Influence Functions for Hessian-Free Data AttributionPhilipp Alexander Kreer, Wilson Wu, Maxwell Adam, Zach Furman et al.ICLR 2026 · 14 citations
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi et al.NeurIPS 2022 · 185 citations
