Make Every Example Count: On the Stability and Utility of Self-Influence for Learning from Noisy NLP Datasets
Irina Bejan, Artem Sokolov, Katja Filippova
Abstract
Increasingly larger datasets have become a standard ingredient to advancing the state-of-the-art in NLP. However, data quality might have already become the bottleneck to unlock further gains. Given the diversity and the sizes of modern datasets, standard data filtering is not straight-forward to apply, because of the multifacetedness of the harmful data and elusiveness of filtering rules that would generalize across multiple tasks. We study the fitness of task-agnostic self-influence scores of training examples for data cleaning, analyze their efficacy in capturing naturally occurring outliers, and investigate to what extent self-influence based data cleaning can improve downstream performance in machine translation, question answering and text classification, building up on recent approaches to self-influence calculation and automated curriculum learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective RolloutsHaizhong Zheng, Yang Zhou, Brian R. Bartoldson, Bhavya Kailkhura et al.NeurIPS 2025 · 125 citations
- β-DPO: Direct Preference Optimization with Dynamic βJunkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu et al.NeurIPS 2024 · 114 citations
- LayerIF: Estimating Layer Quality for Large Language Models using Influence FunctionsHadi Askari, Shivanshu Gupta, Fei Wang, Anshuman Chhabra et al.NeurIPS 2025 · 16 citations
- First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence EstimationDmytro Vitel, Anshuman ChhabraICLR 2026 · 8 citations
- Less is More: High-value Data Selection for Visual Instruction TuningZikang Liu, Kun Zhou, Wayne Xin Zhao, Dawei Gao et al.ACM MM 2025 · 3 citations
Builds on14
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
- SELF: Learning to Filter Noisy Labels with Self-EnsemblingDuc Tam Nguyen, Chaithanya Kumar Mummadi, Thi-Phuong-Nhung Ngo, Thi Hoai Phuong Nguyen et al.ICLR 2020 · 354 citations
- Beyond Synthetic Noise: Deep Learning on Controlled Noisy LabelsLu Jiang, Di Huang, Mason Liu, Weilong YangICML 2020 · 241 citations
Related papers
- Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-TuningJinlong Pang, Na Di, Zhaowei Zhu, Jiaheng Wei et al.ICML 2025
- Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning ModelsAnshuman Chhabra, Bo Li, Jian Chen, Prasant Mohapatra et al.ICML 2025
- G-DIG: Towards Gradient-based DIverse and hiGh-quality Instruction Data Selection for Machine TranslationXingyuan Pan, Luyang Huang, Liyan Kang, Zhicheng Liu et al.ACL 2024
- Reinforced Curriculum Learning on Pre-Trained Neural Machine Translation ModelsMingjun Zhao, Haijiang Wu, Di Niu, Xiaoli WangAAAI 2020 · 46 citations
- Exploring the Efficacy of Automatically Generated Counterfactuals for Sentiment AnalysisLinyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang et al.ACL 2021
