HedgeCut: Maintaining Randomised Trees for Low-Latency Machine Unlearning
Sebastian Schelter, Stefan Grafberger, Ted Dunning
摘要
Software systems that learn from user data with machine learning (ML) have become ubiquitous over the last years. Recent law such as the "General Data Protection Regulation" (GDPR) requires organisations that process personal data to delete user data upon request (enacting the "right to be forgotten"). However, this regulation does not only require the deletion of user data from databases, but also applies to ML models that have been learned from the stored data. We therefore argue that ML applications should offer users to unlearn their data from trained models in a timely manner. We explore how fast this unlearning can be done under the constraints imposed by real world deployments, and introduce the problem of low-latency machine unlearning: maintaining a deployed ML model in-place under the removal of a small fraction of training samples without retraining.
We propose HedgeCut, a classification model based on an ensemble of randomised decision trees, which is designed to answer unlearning requests with low latency. We detail how to efficiently implement HedgeCut with vectorised operators for decision tree learning. We conduct an experimental evaluation on five privacysensitive datasets, where we find that HedgeCut can unlearn training samples with a latency of around 100 microseconds and answers up to 36,000 prediction requests per second, while providing a training time and predictive accuracy similar to widely used implementations of tree-based ML models such as Random Forests.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Machine Unlearning for Random ForestsJonathan Brophy, Daniel LowdICML 2021 · 被引用 222 次
- The Right to be Forgotten in Federated Learning: An Efficient Realization with Rapid RetrainingYi Liu, Lei Xu, Xingliang Yuan, Cong Wang 等INFOCOM 2022 · 被引用 189 次
- Hard to Forget: Poisoning Attacks on Certified Machine UnlearningNeil G. Marchant, Benjamin I. P. Rubinstein, Scott AlfeldAAAI 2022 · 被引用 95 次
- Fast Federated Machine Unlearning with Nonlinear Functional TheoryTianshi Che, Yang Zhou, Zijie Zhang, Lingjuan Lyu 等ICML 2023 · 被引用 77 次
- Prompt Certified Machine Unlearning with Randomized Gradient Smoothing and QuantizationZijie Zhang, Yang Zhou, Xin Zhao, Tianshi Che 等NeurIPS 2022 · 被引用 56 次
它引用的顶会 Paper3
- DeltaGrad: Rapid retraining of machine learning modelsYinjun Wu, Edgar Dobriban, Susan B. DavidsonICML 2020 · 被引用 262 次
- Detecting Violations of Differential PrivacyZeyu Ding, Yuxin Wang, Guanhong Wang, Danfeng Zhang 等CCS 2018 · 被引用 156 次
- Understanding and Benchmarking the Impact of GDPR on Database SystemsSupreeth Shastri, Vinay Banakar, Melissa Wasserman, Arun Kumar 等VLDB 2020 · 被引用 82 次
相关 Paper
- Machine Unlearning in Gradient Boosting Decision TreesHuawei Lin, Jun Woo Chung, Yingjie Lao, Weijie ZhaoKDD 2023 · 被引用 13 次
- DynFrs: An Efficient Framework for Machine Unlearning in Random ForestShurong Wang, Zhuoyang Shen, Xinbao Qiao, Tongning Zhang 等ICLR 2025
- ERASER: Machine Unlearning in MLaaS via an Inference Serving-Aware ApproachYuke Hu, Jian Lou, Jiaqi Liu, Wangze Ni 等CCS 2024 · 被引用 14 次
- DeltaBoost: Gradient Boosting Decision Trees with Efficient Machine UnlearningZhaomin Wu, Junhui Zhu, Qinbin Li, Bingsheng HeSIGMOD 2023 · 被引用 18 次
- Amnesiac Machine LearningLaura Graves, Vineel Nagisetty, Vijay GaneshAAAI 2021 · 被引用 416 次
