Resolving Training Biases via Influence-based Data Relabeling
Shuming Kong, Yanyan Shen, Linpeng Huang
Abstract
The performance of supervised learning methods easily suffers from the training bias issue caused by train-test distribution mismatch or label noise. Influence function is a technique that estimates the impacts of a training sample on the model’s predictions. Recent studies on data resampling have employed influence functions to identify harmful training samples that will degrade model's test performance. They have shown that discarding or downweighting the identified harmful training samples is an effective way to resolve training biases. In this work, we move one step forward and propose an influence-based relabeling framework named RDIA for reusing harmful training samples toward better model performance. To achieve this, we use influence functions to estimate how relabeling a training sample would affect model's test performance and further develop a novel relabeling function R. We theoretically prove that applying R to relabel harmful training samples allows the model to achieve lower test loss than simply discarding them for any classification tasks using cross-entropy loss. Extensive experiments on ten real-world datasets demonstrate RDIA outperforms the state-of-the-art data resampling methods and improves model's robustness against label noise.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7d49bd54-76d3-4e52-86e2-364dc857e7dfCited by top-tier papers36
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc et al.ICML 2023 · 260 citations
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi et al.NeurIPS 2022 · 185 citations
- SoftPatch: Unsupervised Anomaly Detection with Noisy DataXi Jiang, Jianlin Liu, Jinbao Wang, Qiang Nie et al.NeurIPS 2022 · 118 citations
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion ModelsYongchan Kwon, Eric Wu, Kevin Wu, James ZouICLR 2024 · 112 citations
- Intriguing Properties of Data Attribution on Diffusion ModelsXiaosen Zheng, Tianyu Pang, Chao Du, Jing Jiang et al.ICLR 2024 · 41 citations
Related papers
- Data Glitches Discovery using Influence-based Model ExplanationsNikolaos Myrtakis, Ioannis Tsamardinos, Vassilis ChristophidesKDD 2025
- Rescaled Influence Functions: Accurate Data Attribution in High DimensionIttai Rubinstein, Samuel B. HopkinsNeurIPS 2025 · 3 citations
- Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data ValueYongchan Kwon, James ZouICML 2023 · 54 citations
- Influence Estimation for Generative Adversarial NetworksNaoyuki Terashita, Hiroki Ohashi, Yuichi Nonaka, Takashi KanemaruICLR 2021 · 12 citations
- Towards Robust Influence Functions with Flat Validation MinimaXichen Ye, Yifan Wu, Weizhong Zhang, Cheng Jin et al.ICML 2025
