Lune

NeurIPS2021顶会

Deep Learning on a Data Diet: Finding Important Examples Early in Training

Mansheej Paul, Surya Ganguli, Gintare Karolina Dziugaite

2021年份
806被引次数
243顶会引用

摘要

Recent success in deep learning has partially been driven by training increasingly overparametrized networks on ever larger datasets. It is therefore natural to ask: how much of the data is superfluous, which examples are important for generalization, and how do we find them? In this work, we make the striking observation that, in standard vision datasets, simple scores averaged over several weight initializations can be used to identify important examples very early in training. We propose two such scores-the Gradient Normed (GraNd) and the Error L2-Norm (EL2N) scores-and demonstrate their efficacy on a range of architectures and datasets by pruning significant fractions of training data without sacrificing test accuracy. In fact, using EL2N scores calculated a few epochs into training, we can prune half of the CIFAR10 training set while slightly improving test accuracy. Furthermore, for a given dataset, EL2N scores from one architecture or hyperparameter configuration generalize to other configurations. Compared to recent work that prunes data by discarding examples that are rarely forgotten over the course of training, our scores use only local information early in training. We also use our scores to detect noisy examples and study training dynamics through the lens of important examples-we investigate how the data distribution shapes the loss surface and identify subspaces of the model's data representation that are relatively stable over training. * This work was carried out while the author was at ServiceNow. It was finalized at Google Brain. This version supersedes the published NeurIPS 2021 version of this work. Due to a bug in Flax[1] identified by Kirsch [2], the results for the GraNd score computed at initialization were miscalculated, and subsequent conclusions about pruning at initialization were erroneous. This version corrects these errors. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper243

问问它们各自怎么用它

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖