REDUCR: Robust Data Downsampling using Class Priority Reweighting
William Bankes, George Hughes, Ilija Bogunovic, Zi Wang
摘要
Modern machine learning models are becoming increasingly expensive to train for real-world image and text classification tasks, where massive web-scale data is collected in a streaming fashion. To reduce the training cost, online batch selection techniques have been developed to choose the most informative datapoints. However, these techniques can suffer from poor worst-class generalization performance due to class imbalance and distributional shifts. This work introduces REDUCR, a robust and efficient data downsampling method that uses class priority reweighting. REDUCR reduces the training data while preserving worst-class generalization performance. REDUCR assigns priority weights to datapoints in a class-aware manner using an online learning algorithm. We demonstrate the data efficiency and robust performance of REDUCR on vision and text classification tasks. On web-scraped datasets with imbalanced class distributions, REDUCR significantly improves worst-class test accuracy (and average accuracy), surpassing state-of-the-art methods by around 15%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford 等ICLR 2020 · 被引用 974 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
相关 Paper
- Towards Accelerated Model Training via Bayesian Data SelectionZhijie Deng, Peng Cui, Jun ZhuNeurIPS 2023 · 被引用 15 次
- Revive Re-weighting in Imbalanced Learning by Density Ratio EstimationJiaan Luo, Feng Hong, Jiangchao Yao, Bo Han 等NeurIPS 2024 · 被引用 16 次
- Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntSören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma 等ICML 2022 · 被引用 237 次
- FastClass: A Time-Efficient Approach to Weakly-Supervised Text ClassificationTingyu Xia, Yue Wang, Yuan Tian, Yi ChangEMNLP 2022 · 被引用 1 次
- GLISTER: Generalization based Data Subset Selection for Efficient and Robust LearningKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Rishabh K. IyerAAAI 2021 · 被引用 300 次
