Capturing the Temporal Dependence of Training Data Influence
Jiachen T. Wang, Dawn Song, James Zou, Prateek Mittal, Ruoxi Jia
摘要
Traditional data influence estimation methods, like influence function, assume that learning algorithms are permutation-invariant with respect to training data. However, modern training paradigms, especially for foundation models using stochastic algorithms and multi-stage curricula, are sensitive to data ordering, thus violating this assumption. This mismatch renders influence functions inadequate for answering a critical question in machine learning: How can we capture the dependence of data influence on the optimization trajectory during training? To address this gap, we formalize the concept of trajectory-specific leave-one-out (LOO) influence, which quantifies the impact of removing a data point from a specific iteration during training, accounting for the exact sequence of data encountered and the model's optimization trajectory. However, exactly evaluating the trajectory-specific LOO presents a significant computational challenge. To address this, we propose data value embedding, a novel technique enabling efficient approximation of trajectory-specific LOO. Specifically, we compute a training data embedding that encapsulates the cumulative interactions between data and the evolving model parameters. The LOO can then be efficiently approximated through a simple dot-product between the data value embedding and the gradient of the given test data. As data value embedding captures training data ordering, it offers valuable insights into model training dynamics. In particular, we uncover distinct phases of data influence, revealing that data points in the early and late stages of training exert a greater impact on the final model. These insights translate into actionable strategies for managing the computational overhead of data selection by strategically timing the selection process, potentially opening new avenues in data curation research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Group-Level Data Selection for Efficient PretrainingZichun Yu, Fei Peng, Jie Lei, Arnold Overwijk 等NeurIPS 2025 · 被引用 13 次
- A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement LearningYuzheng Hu, Fan Wu, Haotian Ye, David A. Forsyth 等NeurIPS 2025 · 被引用 13 次
- OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every IterationShaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu 等ICML 2026 · 被引用 12 次
- GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse ProjectionPingbang Hu, Joseph Melkonian, Weijing Tang, Han Zhao 等NeurIPS 2025 · 被引用 12 次
- Better Training Data Attribution via Better Inverse Hessian-Vector ProductsAndrew Wang, Elisa Nguyen, Runshi Yang, Juhan Bae 等NeurIPS 2025 · 被引用 12 次
它引用的顶会 Paper17
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc 等ICML 2023 · 被引用 260 次
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi 等NeurIPS 2022 · 被引用 185 次
- Scaling Up Influence FunctionsAndrea Schioppa, Polina Zablotskaia, David Vilar, Artem SokolovAAAI 2022 · 被引用 149 次
相关 Paper
- TRACE: Trajectory-based Activation Change Estimation for Task-specific Data SelectionYe He, Shangzhan Li, Yuxin Zhou, Qi ShiAAAI 2026
- Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data ValueYongchan Kwon, James ZouICML 2023 · 被引用 54 次
- f-INE: A Hypothesis Testing Framework for Estimating Influence under Training RandomnessSubhodip Panda, Dhruv Tarsadiya, Shashwat Sourav, Prathosh AP 等ICLR 2026
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion ModelsYongchan Kwon, Eric Wu, Kevin Wu, James ZouICLR 2024 · 被引用 112 次
- Data Efficient RLVR via Off-Policy Influence GuidanceErle Zhu, Dazhi Jiang, Yuan Wang, Xujun Li 等ACL 2026 · 被引用 7 次
