Learning to Select Pivotal Samples for Meta Re-weighting
Yinjun Wu, Adam Stein, Jacob R. Gardner, Mayur Naik
Abstract
Sample re-weighting strategies provide a promising mechanism to deal with imperfect training data in machine learning, such as noisily labeled or class-imbalanced data. One such strategy involves formulating a bi-level optimization problem called the meta re-weighting problem, whose goal is to optimize performance on a small set of perfect pivotal samples, called meta samples. Many approaches have been proposed to efficiently solve this problem. However, all of them assume that a perfect meta sample set is already provided while we observe that the selections of meta sample set is performance-critical. In this paper, we study how to learn to identify such a meta sample set from a large, imperfect training set, that is subsequently cleaned and used to optimize performance in the meta re-weighting setting. We propose a learning framework which reduces the meta samples selection problem to a weighted K-means clustering problem through rigorously theoretical analysis. We propose two clustering methods within our learning framework, Representation-based clustering method (RBC) and Gradient-based clustering method (GBC), for balancing performance and computational efficiency. Empirical studies demonstrate the performance advantage of our methods over various baseline methods
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b714ee9-319b-4201-8f03-6c16df867a4eBuilds on10
- Contrastive ClusteringYunfan Li, Peng Hu, Jerry Zitao Liu, Dezhong Peng et al.AAAI 2021 · 798 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu et al.ICLR 2022 · 338 citations
- GLISTER: Generalization based Data Subset Selection for Efficient and Robust LearningKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Rishabh K. IyerAAAI 2021 · 300 citations
Related papers
- Learning to Re-weight Examples with Optimal Transport for Imbalanced ClassificationDandan Guo, Zhuo Li, Meixi Zheng, He Zhao et al.NeurIPS 2022 · 46 citations
- Learning Fast Sample Re-weighting Without Reward DataZizhao Zhang, Tomas PfisterICCV 2021 · 109 citations
- Meta-Guided Sample Reweighting for Robust Cross-Modal Hashing Retrieval with Noisy LabelsZiang Tan, Weitao An, Erkun YangAAAI 2026
- Communication-Efficient Robust Federated Learning with Noisy LabelsJunyi Li, Jian Pei, Heng HuangKDD 2022 · 22 citations
- Meta Label Correction for Noisy Label LearningGuoqing Zheng, Ahmed Hassan Awadallah, Susan T. DumaisAAAI 2021 · 239 citations
