Deletion-Anticipative Data Selection with a Limited Budget
Rachael Hwee Ling Sim, Jue Fan, Xiao Tian, Patrick Jaillet, Bryan Kian Hsiang Low
Abstract
Learners with a limited budget can use supervised data subset selection and active learning techniques to select a smaller training set and reduce the cost of acquiring data and training machine learning (ML) models. However, the resulting high model performance, measured by a data utility function, may not be preserved when some data owners, enabled by the GDPR's right to erasure, request their data to be deleted from the ML model. This raises an important question for learners who are temporarily unable or unwilling to acquire data again: During the initial data acquisition of a training set of size k, can we proactively maximize the data utility after future unknown deletions? We propose that the learner anticipates/estimates the probability that (i) each data owner in the feasible set will independently delete its data or (ii) a number of deletions occur out of k, and justify our proposal with concrete real-world use cases. Then, instead of directly maximizing the data utility function, the learner can maximize the expected or risk-averse post-deletion utility based on the anticipated probabilities. We further propose how to construct these deletion-anticipative data selection (DADS) maximization objectives to preserve monotone submodularity and near-optimality of greedy solutions, how to optimize the objectives and empirically evaluate DADS' performance on realworld datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang LowICML 2020 · 158 citations
- Deletion Robust Submodular Maximization over MatroidsPaul Duetting, Federico Fusco, Silvio Lattanzi, Ashkan Norouzi-Fard et al.ICML 2022 · 20 citations
- Training-Free Neural Active Learning with Initialization-Robustness GuaranteesApivich Hemachandra, Zhongxiang Dai, Jasraj Singh, See-Kiong Ng et al.ICML 2023 · 8 citations
Related papers
- DeRDaVa: Deletion-Robust Data Valuation for Machine LearningXiao Tian, Rachael Hwee Ling Sim, Jue Fan, Bryan Kian Hsiang LowAAAI 2024 · 3 citations
- Control, Confidentiality, and the Right to be ForgottenAloni Cohen, Adam D. Smith, Marika Swanberg, Prashant Nalini VasudevanCCS 2023 · 6 citations
- On the Trade-Off between Actionable Explanations and the Right to be ForgottenMartin Pawelczyk, Tobias Leemann, Asia Biega, Gjergji KasneciICLR 2023 · 3 citations
- Machine Unlearning for Random ForestsJonathan Brophy, Daniel LowdICML 2021 · 222 citations
- Amnesiac Machine LearningLaura Graves, Vineel Nagisetty, Vijay GaneshAAAI 2021 · 416 citations
