Trained Random Forests Completely Reveal your Dataset
Julien Ferry, Ricardo Fukasawa, Timothée Pascal, Thibaut Vidal
Abstract
We introduce an optimization-based reconstruction attack capable of completely or near-completely reconstructing a dataset utilized for training a random forest. Notably, our approach relies solely on information readily available in commonly used libraries such as scikit-learn. To achieve this, we formulate the reconstruction problem as a combinatorial problem under a maximum likelihood objective. We demonstrate that this problem is NP-hard, though solvable at scale using constraint programming -- an approach rooted in constraint propagation and solution-domain reduction. Through an extensive computational investigation, we demonstrate that random forests trained without bootstrap aggregation but with feature randomization are susceptible to a complete reconstruction. This holds true even with a small number of trees. Even with bootstrap aggregation, the majority of the data can also be reconstructed. These findings underscore a critical vulnerability inherent in widely adopted ensemble methods, warranting attention and mitigation. Although the potential for such reconstruction attacks has been discussed in privacy research, our study provides clear empirical evidence of their practicability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- From Counterfactuals to Trees: Competitive Analysis of Model Extraction AttacksAwa Khouna, Julien Ferry, Thibaut VidalNeurIPS 2025 · 3 citations
- The Double-Edged Nature of the Rashomon Set for Trustworthy Machine LearningEthan Hsu, Harry Chen, Chudi Zhong, Lesia SemenovaICML 2026 · 1 citation
- RECAST: Model Reconstruction via Counterfactual-Aware Wasserstein Geometry under Limited DataXuan Zhao, Lena Krieger, Zhuo Cao, Arya Bangun et al.ICML 2026
Builds on6
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Reconstructing Training Data From Trained Neural NetworksNiv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir et al.NeurIPS 2022 · 196 citations
- Optimal Counterfactual Explanations in Tree EnsemblesAxel Parmentier, Thibaut VidalICML 2021 · 66 citations
- Optimal Kidney Exchange with ImmunosuppressantsHaris Aziz, Ágnes Cseh, John P. Dickerson, Duncan C. McElfreshAAAI 2021 · 17 citations
Related papers
- On Collective Robustness of Bagging Against Data PoisoningRuoxin Chen, Zenan Li, Jie Li, Junchi Yan et al.ICML 2022 · 25 citations
- Intrinsic Certified Robustness of Bagging against Data Poisoning AttacksJinyuan Jia, Xiaoyu Cao, Neil Zhenqiang GongAAAI 2021 · 155 citations
- Auditing Privacy Mechanisms via Label Inference AttacksRóbert Busa-Fekete, Travis Dick, Claudio Gentile, Andrés Muñoz Medina et al.NeurIPS 2024 · 3 citations
- Verifiable Boosted Tree EnsemblesStefano Calzavara, Lorenzo Cazzaro, Claudio Lucchese, Giulio Ermanno PibiriS&P 2025
- Cost-Aware Robust Tree Ensembles for Security ApplicationsYizheng Chen, Shiqi Wang, Weifan Jiang, Asaf Cidon et al.USENIX Security 2021 · 26 citations
