Machine Unlearning for Random Forests
Jonathan Brophy, Daniel Lowd
Abstract
Responding to user data deletion requests, removing noisy examples, or deleting corrupted training data are just a few reasons for wanting to delete instances from a machine learning (ML) model. However, efficiently removing this data from an ML model is generally difficult. In this paper, we introduce data removal-enabled (DaRE) forests, a variant of random forests that enables the removal of training data with minimal retraining. Model updates for each DaRE tree in the forest are exact, meaning that removing instances from a DaRE model yields exactly the same model as retraining from scratch on updated data. DaRE trees use randomness and caching to make data deletion efficient. The upper levels of DaRE trees use random nodes, which choose split attributes and thresholds uniformly at random. These nodes rarely require updates because they only minimally depend on the data. At the lower levels, splits are chosen to greedily optimize a split criterion such as Gini index or mutual information. DaRE trees cache statistics at each node and training data at each leaf, so that only the necessary subtrees are updated as data is removed. For numerical attributes, greedy nodes optimize over a random subset of thresholds, so that they can maintain statistics while approximating the optimal threshold. By adjusting the number of thresholds considered for greedy nodes, and the number of random nodes, DaRE trees can trade off between more accurate predictions and more efficient updates. In experiments on 13 real-world datasets and one synthetic dataset, we find DaRE forests delete data orders of magnitude faster than retraining from scratch while sacrificing little to no predictive power.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 112aa331-509b-42d7-9a10-7a34d1d43ad0Cited by top-tier papers62
- Can Bad Teaching Induce Forgetting? Unlearning in Deep Networks Using an Incompetent TeacherVikram S. Chundawat, Ayush K. Tarun, Murari Mandal, Mohan S. KankanhalliAAAI 2023 · 247 citations
- Federated Unlearning via Class-Discriminative PruningJunxiao Wang, Song Guo, Xin Xie, Heng QiWWW 2022 · 217 citations
- The Right to be Forgotten in Federated Learning: An Efficient Realization with Rapid RetrainingYi Liu, Lei Xu, Xingliang Yuan, Cong Wang et al.INFOCOM 2022 · 189 citations
- Fast Model DeBias with Machine UnlearningRuizhe Chen, Jianfei Yang, Huimin Xiong, Jianhong Bai et al.NeurIPS 2023 · 110 citations
- Graph UnlearningMin Chen, Zhikun Zhang, Tianhao Wang, Michael Backes et al.CCS 2022 · 103 citations
Builds on14
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
Related papers
- DeltaBoost: Gradient Boosting Decision Trees with Efficient Machine UnlearningZhaomin Wu, Junhui Zhu, Qinbin Li, Bingsheng HeSIGMOD 2023 · 18 citations
- DeltaGrad: Rapid retraining of machine learning modelsYinjun Wu, Edgar Dobriban, Susan B. DavidsonICML 2020 · 262 citations
- DynFrs: An Efficient Framework for Machine Unlearning in Random ForestShurong Wang, Zhuoyang Shen, Xinbao Qiao, Tongning Zhang et al.ICLR 2025
- Machine Unlearning in Gradient Boosting Decision TreesHuawei Lin, Jun Woo Chung, Yingjie Lao, Weijie ZhaoKDD 2023 · 13 citations
- Factor Decorrelation Enhanced Data Removal from Deep Predictive ModelsWenhao Yang, Lin Li, Xiaohui Tao, Kaize ShiNeurIPS 2025
