Influence Dynamics and Stagewise Data Attribution
Jin Hwa Lee, Matthew Smith, Maxwell Adam, Jesse Hoogland
摘要
Current training data attribution (TDA) methods treat the influence one sample has on another as static, but neural networks learn in distinct stages that exhibit changing patterns of influence. In this work, we introduce a framework for stagewise data attribution grounded in singular learning theory. We predict that influence can change non-monotonically, including sign flips and sharp peaks at developmental transitions. We first validate these predictions analytically and empirically in a toy model, showing that dynamic shifts in influence directly map to the model's progressive learning of a semantic hierarchy. Finally, we demonstrate these phenomena at scale in language models, where token-level influence changes align with known developmental stages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- LLM Circuit Analyses Are Consistent Across Training and ScaleCurt Tigges, Michael Hanna, Qinan Yu, Stella BidermanNeurIPS 2024 · 被引用 71 次
- HyDRA: Hypergradient Data Relevance Analysis for Interpreting Deep Neural NetworksYuanyuan Chen, Boyang Li, Han Yu, Pengcheng Wu 等AAAI 2021 · 被引用 50 次
- Training Data Attribution via Approximate UnrollingJuhan Bae, Wu Lin, Jonathan Lorraine, Roger B. GrosseNeurIPS 2024 · 被引用 41 次
相关 Paper
- Distributional Training Data Attribution: What do Influence Functions Sample?Bruno Kacper Mlodozeniec, Isaac Reid, Sam Power, David Krueger 等NeurIPS 2025
- Scalable Influence and Fact Tracing for Large Language Model PretrainingTyler A. Chang, Dheeraj Rajagopal, Tolga Bolukbasi, Lucas Dixon 等ICLR 2025 · 被引用 1 次
- Who Gets Credit or Blame? Attributing Accountability in Modern AI SystemsShichang Zhang, Hongzhe Du, Jiaqi Ma, Himabindu LakkarajuICML 2026
- Differentiation and Specialization of Attention Heads via the Refined Local Learning CoefficientGeorge Wang, Jesse Hoogland, Stan van Wingerden, Zach Furman 等ICLR 2025
- Enhancing Training Data Attribution with Representational OptimizationWeiwei Sun, Haokun Liu, Nikhil Kandpal, Colin A. Raffel 等NeurIPS 2025 · 被引用 9 次
