Lune

ICLR2024顶会

Towards a statistical theory of data selection under weak supervision

Germain Kolossov, Andrea Montanari, Pulkit Tandon

2024年份
27被引次数
21顶会引用

摘要

Given a sample of size NN, it is often useful to select a subsample of smaller size n<Nn<N to be used for statistical estimation or learning. Such a data selection step is useful to reduce the requirements of data labeling and the computational complexity of learning. We assume to be given NN unlabeled samples {xi}i≤N\{{\boldsymbol x}_i\}_{i\le N}, and to be given access to a `surrogate model' that can predict labels yiy_i better than random guessing. Our goal is to select a subset of the samples, to be denoted by {xi}i∈G\{{\boldsymbol x}_i\}_{i\in G}, of size ∣G∣=n<N|G|=n<N. We then acquire labels for this set and we use them to train a model via regularized empirical risk minimization. By using a mixture of numerical experiments on real and synthetic data, and mathematical derivations under low- and high- dimensional asymptotics, we show that: (i)(i) Data selection can be very effective, in particular beating training on the full sample in some cases; (ii)(ii) Certain popular choices in data selection methods (e.g. unbiased reweighted subsampling, or influence function-based subsampling) can be substantially suboptimal.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper21

问问它们各自怎么用它

它引用的顶会 Paper3

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖