DeRDaVa: Deletion-Robust Data Valuation for Machine Learning
Xiao Tian, Rachael Hwee Ling Sim, Jue Fan, Bryan Kian Hsiang Low
Abstract
Data valuation is concerned with determining a fair valuation of data from data sources to compensate them or to identify training examples that are the most or least useful for predictions. With the rising interest in personal data ownership and data protection regulations, model owners will likely have to fulfil more data deletion requests. This raises issues that have not been addressed by existing works: Are the data valuation scores still fair with deletions? Must the scores be expensively recomputed? The answer is no. To avoid recomputations, we propose using our data valuation framework DeRDaVa upfront for valuing each data source's contribution to preserving robust model performance after anticipated data deletions. DeRDaVa can be efficiently approximated and will assign higher values to data that are more useful or less likely to be deleted. We further generalize DeRDaVa to Risk-DeRDaVa to cater to risk-averse/seeking model owners who are concerned with the worst/best-cases model utility. We also empirically demonstrate the practicality of our solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Is Data Shapley Not Better than Random in Data Selection? Ask NASHXiao Tian, Jue Fan, Rachael Hwee Ling Sim, Zixuan Wang et al.ICML 2026
- INO-SGD: Addressing Utility Imbalance under Individualized Differential PrivacyXiao Tian, Jue Fan, Rachael Hwee Ling Sim, Bryan Kian Hsiang LowICLR 2026
Builds on10
- Adaptive Machine UnlearningVarun Gupta, Christopher Jung, Seth Neel, Aaron Roth et al.NeurIPS 2021 · 262 citations
- Unlearnable Examples: Making Personal Data UnexploitableHanxun Huang, Xingjun Ma, Sarah Monazam Erfani, James Bailey et al.ICLR 2021 · 255 citations
- Variational Bayesian UnlearningQuoc Phong Nguyen, Bryan Kian Hsiang Low, Patrick JailletNeurIPS 2020 · 198 citations
- Incentive Mechanism for Horizontal Federated Learning Based on Reputation and Reverse AuctionJingwen Zhang, Yuezhou Wu, Rong PanWWW 2021 · 176 citations
- Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang LowICML 2020 · 158 citations
Related papers
- Deletion-Anticipative Data Selection with a Limited BudgetRachael Hwee Ling Sim, Jue Fan, Xiao Tian, Patrick Jaillet et al.ICML 2024 · 1 citation
- Distributionally Robust Data ValuationXiaoqiang Lin, Xinyi Xu, Zhaoxuan Wu, See-Kiong Ng et al.ICML 2024 · 6 citations
- Factor Decorrelation Enhanced Data Removal from Deep Predictive ModelsWenhao Yang, Lin Li, Xiaohui Tao, Kaize ShiNeurIPS 2025
- Towards Bridging the Gaps between the Right to Explanation and the Right to be ForgottenSatyapriya Krishna, Jiaqi Ma, Himabindu LakkarajuICML 2023 · 14 citations
- On the Trade-Off between Actionable Explanations and the Right to be ForgottenMartin Pawelczyk, Tobias Leemann, Asia Biega, Gjergji KasneciICLR 2023 · 3 citations
