Influence Dynamics and Stagewise Data Attribution
Jin Hwa Lee, Matthew Smith, Maxwell Adam, Jesse Hoogland
Abstract
Current training data attribution (TDA) methods treat the influence one sample has on another as static, but neural networks learn in distinct stages that exhibit changing patterns of influence. In this work, we introduce a framework for stagewise data attribution grounded in singular learning theory. We predict that influence can change non-monotonically, including sign flips and sharp peaks at developmental transitions. We first validate these predictions analytically and empirically in a toy model, showing that dynamic shifts in influence directly map to the model's progressive learning of a semantic hierarchy. Finally, we demonstrate these phenomena at scale in language models, where token-level influence changes align with known developmental stages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbdfbdfc-50c9-40cc-9492-eed35d09e8a5Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- LLM Circuit Analyses Are Consistent Across Training and ScaleCurt Tigges, Michael Hanna, Qinan Yu, Stella BidermanNeurIPS 2024 · 71 citations
- HyDRA: Hypergradient Data Relevance Analysis for Interpreting Deep Neural NetworksYuanyuan Chen, Boyang Li, Han Yu, Pengcheng Wu et al.AAAI 2021 · 50 citations
- Training Data Attribution via Approximate UnrollingJuhan Bae, Wu Lin, Jonathan Lorraine, Roger B. GrosseNeurIPS 2024 · 41 citations
Related papers
- Distributional Training Data Attribution: What do Influence Functions Sample?Bruno Kacper Mlodozeniec, Isaac Reid, Sam Power, David Krueger et al.NeurIPS 2025
- Scalable Influence and Fact Tracing for Large Language Model PretrainingTyler A. Chang, Dheeraj Rajagopal, Tolga Bolukbasi, Lucas Dixon et al.ICLR 2025 · 1 citation
- Who Gets Credit or Blame? Attributing Accountability in Modern AI SystemsShichang Zhang, Hongzhe Du, Jiaqi Ma, Himabindu LakkarajuICML 2026
- Differentiation and Specialization of Attention Heads via the Refined Local Learning CoefficientGeorge Wang, Jesse Hoogland, Stan van Wingerden, Zach Furman et al.ICLR 2025
- Enhancing Training Data Attribution with Representational OptimizationWeiwei Sun, Haokun Liu, Nikhil Kandpal, Colin A. Raffel et al.NeurIPS 2025 · 9 citations
