Training Data Attribution via Approximate Unrolling
Juhan Bae, Wu Lin, Jonathan Lorraine, Roger B. Grosse
Abstract
Many training data attribution (TDA) methods aim to estimate how a model’s behavior would change if one or more data points were removed from the training set. Methods based on implicit differentiation, such as influence functions, can be made computationally efficient, but fail to account for underspecification, the implicit bias of the optimization algorithm, or multi-stage training pipelines. By contrast, methods based on unrolling address these issues but face scalability challenges. In this work, we connect the implicit-differentiation-based and unrolling-based approaches and combine their benefits by introducing S OURCE , an approximate unrolling-based TDA method that is computed using an influence-function-like formula. While being computationally efficient compared to unrolling-based approaches, S OURCE is suitable in cases where implicit-differentiation-based approaches struggle, such as in non-converged models and multi-stage training pipelines. Empirically, S OURCE outperforms existing TDA techniques in coun-terfactual prediction, especially in settings where implicit-differentiation-based approaches fall short.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ebdfec2-bdf3-45a4-8d9b-6571798e09cbCited by top-tier papers18
- Bayesian Influence Functions for Hessian-Free Data AttributionPhilipp Alexander Kreer, Wilson Wu, Maxwell Adam, Zach Furman et al.ICLR 2026 · 14 citations
- Influence-Preserving Proxies for Gradient-Based Data Selection in LLM FineTuningSirui Chen, Yunzhe Qi, Mengting Ai, Yifan Sun et al.ICLR 2026 · 9 citations
- First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence EstimationDmytro Vitel, Anshuman ChhabraICLR 2026 · 8 citations
- Fast Data Attribution for Text-to-Image ModelsSheng-Yu Wang, Aaron Hertzmann, Alexei A. Efros, Richard Zhang et al.NeurIPS 2025 · 7 citations
- Scalable Valuation of Human Feedback through Provably Robust Model AlignmentMasahiro Fujisawa, Masaki Adachi, Michael A. OsborneNeurIPS 2025 · 6 citations
Builds on30
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
Related papers
- Better Training Data Attribution via Better Inverse Hessian-Vector ProductsAndrew Wang, Elisa Nguyen, Runshi Yang, Juhan Bae et al.NeurIPS 2025 · 12 citations
- Final-Model-Only Data Attribution with a Unifying View of Gradient-Based MethodsDennis Wei, Inkit Padhi, Soumya Ghosh, Amit Dhurandhar et al.NeurIPS 2025 · 6 citations
- Distributional Training Data Attribution: What do Influence Functions Sample?Bruno Kacper Mlodozeniec, Isaac Reid, Sam Power, David Krueger et al.NeurIPS 2025
- Enhancing Training Data Attribution with Representational OptimizationWeiwei Sun, Haokun Liu, Nikhil Kandpal, Colin A. Raffel et al.NeurIPS 2025 · 9 citations
- Influence Functions for Scalable Data Attribution in Diffusion ModelsBruno Kacper Mlodozeniec, Runa Eschenhagen, Juhan Bae, Alexander Immer et al.ICLR 2025
