Leveraging Importance Weights in Subset Selection
Gui Citovsky, Giulia DeSalvo, Sanjiv Kumar, Srikumar Ramalingam, Afshin Rostamizadeh, Yunjuan Wang
Abstract
We present a subset selection algorithm designed to work with arbitrary model families in a practical batch setting. In such a setting, an algorithm can sample examples one at a time but, in order to limit overhead costs, is only able to update its state (i.e. further train model weights) once a large enough batch of examples is selected. Our algorithm, IWeS, selects examples by importance sampling where the sampling probability assigned to each example is based on the entropy of models trained on previously selected batches. IWeS admits significant performance improvement compared to other subset selection algorithms for seven publicly available datasets. Additionally, it is competitive in an active learning setting, where the label information is not available at selection time. We also provide an initial theoretical analysis to support our importance weighting approach, proving generalization and sampling rate bounds.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- BWS: Best Window Selection Based on Sample Scores for Data Pruning across Broad RangesHoyong Choi, Nohyun Ki, Hye Won ChungICML 2024 · 9 citations
- Patch-Aware Sample Selection for Efficient Masked Image ModelingZhengyang Zhuge, Jiaxing Wang, Yong Li, Yongjun Bao et al.AAAI 2024 · 4 citations
Builds on6
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model TrainingKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De et al.ICML 2021 · 305 citations
- Batch Active Learning at ScaleGui Citovsky, Giulia DeSalvo, Claudio Gentile, Lazaros Karydas et al.NeurIPS 2021 · 220 citations
Related papers
- Diversified Batch Selection for Training AccelerationFeng Hong, Yueming Lyu, Jiangchao Yao, Ya Zhang et al.ICML 2024 · 16 citations
- Influence Selection for Active LearningZhuoming Liu, Hao Ding, Huaping Zhong, Weijia Li et al.ICCV 2021 · 125 citations
- Structural-Entropy-Based Sample Selection for Efficient and Effective LearningTianchi Xie, Jiangning Zhu, Guozu Ma, Minzhi Lin et al.ICLR 2025
- Generator Assisted Mixture of Experts for Feature Acquisition in BatchVedang Asgaonkar, Aditya Jain, Abir DeAAAI 2024 · 3 citations
- Entropic Open-Set Active LearningBardia Safaei, Vibashan VS, Celso M. de Melo, Vishal M. PatelAAAI 2024 · 36 citations
