Deletion-Anticipative Data Selection with a Limited Budget
Rachael Hwee Ling Sim, Jue Fan, Xiao Tian, Patrick Jaillet, Bryan Kian Hsiang Low
摘要
Learners with a limited budget can use supervised data subset selection and active learning techniques to select a smaller training set and reduce the cost of acquiring data and training machine learning (ML) models. However, the resulting high model performance, measured by a data utility function, may not be preserved when some data owners, enabled by the GDPR's right to erasure, request their data to be deleted from the ML model. This raises an important question for learners who are temporarily unable or unwilling to acquire data again: During the initial data acquisition of a training set of size k, can we proactively maximize the data utility after future unknown deletions? We propose that the learner anticipates/estimates the probability that (i) each data owner in the feasible set will independently delete its data or (ii) a number of deletions occur out of k, and justify our proposal with concrete real-world use cases. Then, instead of directly maximizing the data utility function, the learner can maximize the expected or risk-averse post-deletion utility based on the anticipated probabilities. We further propose how to construct these deletion-anticipative data selection (DADS) maximization objectives to preserve monotone submodularity and near-optimality of greedy solutions, how to optimize the objectives and empirically evaluate DADS' performance on realworld datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang LowICML 2020 · 被引用 158 次
- Deletion Robust Submodular Maximization over MatroidsPaul Duetting, Federico Fusco, Silvio Lattanzi, Ashkan Norouzi-Fard 等ICML 2022 · 被引用 20 次
- Training-Free Neural Active Learning with Initialization-Robustness GuaranteesApivich Hemachandra, Zhongxiang Dai, Jasraj Singh, See-Kiong Ng 等ICML 2023 · 被引用 8 次
相关 Paper
- DeRDaVa: Deletion-Robust Data Valuation for Machine LearningXiao Tian, Rachael Hwee Ling Sim, Jue Fan, Bryan Kian Hsiang LowAAAI 2024 · 被引用 3 次
- Control, Confidentiality, and the Right to be ForgottenAloni Cohen, Adam D. Smith, Marika Swanberg, Prashant Nalini VasudevanCCS 2023 · 被引用 6 次
- On the Trade-Off between Actionable Explanations and the Right to be ForgottenMartin Pawelczyk, Tobias Leemann, Asia Biega, Gjergji KasneciICLR 2023 · 被引用 3 次
- Machine Unlearning for Random ForestsJonathan Brophy, Daniel LowdICML 2021 · 被引用 222 次
- Amnesiac Machine LearningLaura Graves, Vineel Nagisetty, Vijay GaneshAAAI 2021 · 被引用 416 次
