Trained Random Forests Completely Reveal your Dataset
Julien Ferry, Ricardo Fukasawa, Timothée Pascal, Thibaut Vidal
摘要
We introduce an optimization-based reconstruction attack capable of completely or near-completely reconstructing a dataset utilized for training a random forest. Notably, our approach relies solely on information readily available in commonly used libraries such as scikit-learn. To achieve this, we formulate the reconstruction problem as a combinatorial problem under a maximum likelihood objective. We demonstrate that this problem is NP-hard, though solvable at scale using constraint programming -- an approach rooted in constraint propagation and solution-domain reduction. Through an extensive computational investigation, we demonstrate that random forests trained without bootstrap aggregation but with feature randomization are susceptible to a complete reconstruction. This holds true even with a small number of trees. Even with bootstrap aggregation, the majority of the data can also be reconstructed. These findings underscore a critical vulnerability inherent in widely adopted ensemble methods, warranting attention and mitigation. Although the potential for such reconstruction attacks has been discussed in privacy research, our study provides clear empirical evidence of their practicability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- From Counterfactuals to Trees: Competitive Analysis of Model Extraction AttacksAwa Khouna, Julien Ferry, Thibaut VidalNeurIPS 2025 · 被引用 3 次
- The Double-Edged Nature of the Rashomon Set for Trustworthy Machine LearningEthan Hsu, Harry Chen, Chudi Zhong, Lesia SemenovaICML 2026 · 被引用 1 次
- RECAST: Model Reconstruction via Counterfactual-Aware Wasserstein Geometry under Limited DataXuan Zhao, Lena Krieger, Zhuo Cao, Arya Bangun 等ICML 2026
它引用的顶会 Paper6
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song 等S&P 2022 · 被引用 1,049 次
- Reconstructing Training Data From Trained Neural NetworksNiv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir 等NeurIPS 2022 · 被引用 196 次
- Optimal Counterfactual Explanations in Tree EnsemblesAxel Parmentier, Thibaut VidalICML 2021 · 被引用 66 次
- Optimal Kidney Exchange with ImmunosuppressantsHaris Aziz, Ágnes Cseh, John P. Dickerson, Duncan C. McElfreshAAAI 2021 · 被引用 17 次
相关 Paper
- On Collective Robustness of Bagging Against Data PoisoningRuoxin Chen, Zenan Li, Jie Li, Junchi Yan 等ICML 2022 · 被引用 25 次
- Intrinsic Certified Robustness of Bagging against Data Poisoning AttacksJinyuan Jia, Xiaoyu Cao, Neil Zhenqiang GongAAAI 2021 · 被引用 155 次
- Auditing Privacy Mechanisms via Label Inference AttacksRóbert Busa-Fekete, Travis Dick, Claudio Gentile, Andrés Muñoz Medina 等NeurIPS 2024 · 被引用 3 次
- Verifiable Boosted Tree EnsemblesStefano Calzavara, Lorenzo Cazzaro, Claudio Lucchese, Giulio Ermanno PibiriS&P 2025
- Cost-Aware Robust Tree Ensembles for Security ApplicationsYizheng Chen, Shiqi Wang, Weifan Jiang, Asaf Cidon 等USENIX Security 2021 · 被引用 26 次
