Machine Unlearning for Random Forests
Jonathan Brophy, Daniel Lowd
摘要
Responding to user data deletion requests, removing noisy examples, or deleting corrupted training data are just a few reasons for wanting to delete instances from a machine learning (ML) model. However, efficiently removing this data from an ML model is generally difficult. In this paper, we introduce data removal-enabled (DaRE) forests, a variant of random forests that enables the removal of training data with minimal retraining. Model updates for each DaRE tree in the forest are exact, meaning that removing instances from a DaRE model yields exactly the same model as retraining from scratch on updated data. DaRE trees use randomness and caching to make data deletion efficient. The upper levels of DaRE trees use random nodes, which choose split attributes and thresholds uniformly at random. These nodes rarely require updates because they only minimally depend on the data. At the lower levels, splits are chosen to greedily optimize a split criterion such as Gini index or mutual information. DaRE trees cache statistics at each node and training data at each leaf, so that only the necessary subtrees are updated as data is removed. For numerical attributes, greedy nodes optimize over a random subset of thresholds, so that they can maintain statistics while approximating the optimal threshold. By adjusting the number of thresholds considered for greedy nodes, and the number of random nodes, DaRE trees can trade off between more accurate predictions and more efficient updates. In experiments on 13 real-world datasets and one synthetic dataset, we find DaRE forests delete data orders of magnitude faster than retraining from scratch while sacrificing little to no predictive power.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper62
- Can Bad Teaching Induce Forgetting? Unlearning in Deep Networks Using an Incompetent TeacherVikram S. Chundawat, Ayush K. Tarun, Murari Mandal, Mohan S. KankanhalliAAAI 2023 · 被引用 247 次
- Federated Unlearning via Class-Discriminative PruningJunxiao Wang, Song Guo, Xin Xie, Heng QiWWW 2022 · 被引用 217 次
- The Right to be Forgotten in Federated Learning: An Efficient Realization with Rapid RetrainingYi Liu, Lei Xu, Xingliang Yuan, Cong Wang 等INFOCOM 2022 · 被引用 189 次
- Fast Model DeBias with Machine UnlearningRuizhe Chen, Jianfei Yang, Huimin Xiong, Jianhong Bai 等NeurIPS 2023 · 被引用 110 次
- Graph UnlearningMin Chen, Zhikun Zhang, Tianhao Wang, Michael Backes 等CCS 2022 · 被引用 103 次
它引用的顶会 Paper14
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
相关 Paper
- DeltaBoost: Gradient Boosting Decision Trees with Efficient Machine UnlearningZhaomin Wu, Junhui Zhu, Qinbin Li, Bingsheng HeSIGMOD 2023 · 被引用 18 次
- DeltaGrad: Rapid retraining of machine learning modelsYinjun Wu, Edgar Dobriban, Susan B. DavidsonICML 2020 · 被引用 262 次
- DynFrs: An Efficient Framework for Machine Unlearning in Random ForestShurong Wang, Zhuoyang Shen, Xinbao Qiao, Tongning Zhang 等ICLR 2025
- Machine Unlearning in Gradient Boosting Decision TreesHuawei Lin, Jun Woo Chung, Yingjie Lao, Weijie ZhaoKDD 2023 · 被引用 13 次
- Factor Decorrelation Enhanced Data Removal from Deep Predictive ModelsWenhao Yang, Lin Li, Xiaohui Tao, Kaize ShiNeurIPS 2025
