Lune

NeurIPS2024Top-tier venue

Training Data Attribution via Approximate Unrolling

Juhan Bae, Wu Lin, Jonathan Lorraine, Roger B. Grosse

2024Year
41Citations
18Top-tier citations

Abstract

Many training data attribution (TDA) methods aim to estimate how a model’s behavior would change if one or more data points were removed from the training set. Methods based on implicit differentiation, such as influence functions, can be made computationally efficient, but fail to account for underspecification, the implicit bias of the optimization algorithm, or multi-stage training pipelines. By contrast, methods based on unrolling address these issues but face scalability challenges. In this work, we connect the implicit-differentiation-based and unrolling-based approaches and combine their benefits by introducing S OURCE , an approximate unrolling-based TDA method that is computed using an influence-function-like formula. While being computationally efficient compared to unrolling-based approaches, S OURCE is suitable in cases where implicit-differentiation-based approaches struggle, such as in non-converged models and multi-stage training pipelines. Empirically, S OURCE outperforms existing TDA techniques in coun-terfactual prediction, especially in settings where implicit-differentiation-based approaches fall short.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9ebdfec2-bdf3-45a4-8d9b-6571798e09cb

Cited by top-tier papers18

Ask how each one uses it

Builds on30

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines