First is Better Than Last for Language Data Influence
Chih-Kuan Yeh, Ankur Taly, Mukund Sundararajan, Frederick Liu, Pradeep Ravikumar
摘要
The ability to identify influential training examples enables us to debug training data and explain model behavior. Existing techniques to do so are based on the flow of training data influence through the model parameters (Koh & Liang, 2017; Yeh et al., 2018; Pruthi et al., 2020) . For large models in NLP applications, it is often computationally infeasible to study this flow through all model parameters, therefore techniques usually pick the last layer of weights. However, we observe that since the activation connected to the last layer of weights contains "shared logic", the data influenced calculated via the last layer weights prone to a "cancellation effect", where the data influence of different examples have large magnitude that contradicts each other. The cancellation effect lowers the discriminative power of the influence score, and deleting influential examples according to this measure often does not change the model's behavior by much. To mitigate this, we propose a technique called TracIn-WE that modifies a method called TracIn (Pruthi et al., 2020) to operate on the word embedding layer instead of the last layer, where the cancellation effect is less severe. One potential concern is that influence based on the word embedding layer may not encode sufficient high level information. However, we find that gradients (unlike embeddings) do not suffer from this, possibly because they chain through higher layers. We show that TracIn-WE significantly outperforms other data influence methods applied on the last layer significantly on the case deletion evaluation on three language classification tasks for different models. In addition, TracIn-WE can produce scores not just at the level of the overall training input, but also at the level of words within the training input, a further aid in debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsSang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao 等NeurIPS 2025 · 被引用 112 次
- DOGE: Domain Reweighting with Generalization EstimationSimin Fan, Matteo Pagliardini, Martin JaggiICML 2024 · 被引用 79 次
- Massive Editing for Large Language Models via Meta LearningChenmien Tan, Ge Zhang, Jie FuICLR 2024 · 被引用 68 次
- Training Data Attribution via Approximate UnrollingJuhan Bae, Wu Lin, Jonathan Lorraine, Roger B. GrosseNeurIPS 2024 · 被引用 41 次
- Helpful or Harmful Data? Fine-tuning-free Shapley Attribution for Explaining Language Model PredictionsJingtan Wang, Xiaoqiang Lin, Rui Qiao, Chuan-Sheng Foo 等ICML 2024 · 被引用 12 次
它引用的顶会 Paper6
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence FunctionsXiaochuang Han, Byron C. Wallace, Yulia TsvetkovACL 2020 · 被引用 91 次
- Evaluation of Similarity-based ExplanationsKazuaki Hanawa, Sho Yokoi, Satoshi Hara, Kentaro InuiICLR 2021 · 被引用 79 次
- Representer Point Selection via Local Jacobian Expansion for Post-hoc Classifier Explanation of Deep Neural Networks and Ensemble ModelsYi Sui, Ga Wu, Scott SannerNeurIPS 2021 · 被引用 26 次
相关 Paper
- First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence EstimationDmytro Vitel, Anshuman ChhabraICLR 2026 · 被引用 8 次
- How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuningRochelle Choenni, Dan Garrette, Ekaterina ShutovaEMNLP 2023 · 被引用 2 次
- Scalable Influence and Fact Tracing for Large Language Model PretrainingTyler A. Chang, Dheeraj Rajagopal, Tolga Bolukbasi, Lucas Dixon 等ICLR 2025 · 被引用 1 次
- Influence Scores at Scale for Efficient Language Data SamplingNikhil Anand, Joshua Tan, Maria MinakovaEMNLP 2023 · 被引用 1 次
- Understanding Influence Functions and Datamodels via Harmonic AnalysisNikunj Saunshi, Arushi Gupta, Mark Braverman, Sanjeev AroraICLR 2023 · 被引用 1 次
