Influence Scores at Scale for Efficient Language Data Sampling
Nikhil Anand, Joshua Tan, Maria Minakova
Abstract
Modern ML systems ingest data aggregated from diverse sources, such as synthetic, human-annotated, and live customer traffic. Understanding which examples are important to the performance of a learning algorithm is crucial for efficient model training. Recently, a growing body of literature has given rise to various “influence scores,” which use training artifacts such as model confidence or checkpointed gradients to identify important subsets of data. However, these methods have primarily been developed in computer vision settings, and it remains unclear how well they generalize to language-based tasks using pretrained models. In this paper, we explore the applicability of influence scores in language classification tasks. We evaluate a diverse subset of these scores on the SNLI dataset by quantifying accuracy changes in response to pruning training data through random and influence-score-based sampling. We then stress-test one of the scores – “variance of gradients” (VoG) from Agarwal and Hooker (2022) – in an NLU model stack that was exposed to dynamic user speech patterns in a voice assistant type of setting. Our experiments demonstrate that in many cases, encoder-based language models can be fine-tuned on roughly 50% of the original data without degradation in performance metrics. Along the way, we summarize lessons learned from applying out-of-the-box implementations of influence scores, quantify the effects of noisy and class-imbalanced data, and offer recommendations on score-based sampling for better accuracy and training efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11859233-08f5-4722-b98c-ce39ae76009aCited by top-tier papers4
- TarDiff: Target-Oriented Diffusion Guidance for Synthetic Electronic Health Record Time Series GenerationBowen Deng, Chang Xu, Hao Li, Yu-Hao Huang et al.KDD 2025
- RepLLM: Toward Automatically Reproducing Network Research ResultsYining Jiang, Yunxin Xu, Wenyun Xu, Yufan Zhu et al.SIGCOMM 2026
- Enhancing Training Data Attribution for Large Language Models with Fitting Error ConsiderationKangxi Wu, Liang Pang, Huawei Shen, Xueqi ChengEMNLP 2024
- Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data SelectionJianghao Wu, Jianfei Cai, Weiqiang Wang, Jin Ye et al.ICML 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
Related papers
- First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence EstimationDmytro Vitel, Anshuman ChhabraICLR 2026 · 8 citations
- Estimating Example Difficulty using Variance of GradientsChirag Agarwal, Daniel D'souza, Sara HookerCVPR 2022 · 57 citations
- Self-Influence Guided Data Reweighting for Language Model Pre-trainingMegh Thakkar, Tolga Bolukbasi, Sriram Ganapathy, Shikhar Vashishth et al.EMNLP 2023 · 2 citations
- Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning ModelsAnshuman Chhabra, Bo Li, Jian Chen, Prasant Mohapatra et al.ICML 2025
- Efficient Data Selection at Scale via Influence DistillationMahdi Nikdan, Vincent Cohen-Addad, Dan Alistarh, Vahab MirrokniNeurIPS 2025 · 15 citations
