Lune

ICLR2022Top-tier venue

Resolving Training Biases via Influence-based Data Relabeling

Shuming Kong, Yanyan Shen, Linpeng Huang

2022Year
71Citations
36Top-tier citations

Abstract

The performance of supervised learning methods easily suffers from the training bias issue caused by train-test distribution mismatch or label noise. Influence function is a technique that estimates the impacts of a training sample on the model’s predictions. Recent studies on data resampling have employed influence functions to identify harmful training samples that will degrade model's test performance. They have shown that discarding or downweighting the identified harmful training samples is an effective way to resolve training biases. In this work, we move one step forward and propose an influence-based relabeling framework named RDIA for reusing harmful training samples toward better model performance. To achieve this, we use influence functions to estimate how relabeling a training sample would affect model's test performance and further develop a novel relabeling function R. We theoretically prove that applying R to relabel harmful training samples allows the model to achieve lower test loss than simply discarding them for any classification tasks using cross-entropy loss. Extensive experiments on ten real-world datasets demonstrate RDIA outperforms the state-of-the-art data resampling methods and improves model's robustness against label noise.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 7d49bd54-76d3-4e52-86e2-364dc857e7df

Cited by top-tier papers36

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines