ReTrace: Reinforcement Learning-Guided Reconstruction Attacks on Machine Unlearning
Mengyao Ma, Shuofeng Liu, Minhui Xue, Surya Nepal, Guangdong Bai
Abstract
Machine unlearning has emerged as an inevitable AI mechanism to support GDPR requirements such as revoking user consent through the "right to be forgotten". However, existing approaches often leave residual traces that make them vulnerable to data reconstruction attacks. In this work, we propose ReTrace, the first reconstruction attack framework that uniquely formulates unlearned data recovery on large-scale deep architectures as a reinforcement learning (RL) problem. By treating residual unlearning traces as reward signals, ReTrace guides a generator to actively explore the input space and converge toward the forgotten data distribution. This RL-guided approach enables both instance-level recovery of individual samples and distribution-level reconstruction of unlearned classes. We provide a theoretical foundation showing that the RL objective converges to an exponential-tilted distribution that amplifies forgotten regions. Empirically, ReTrace achieves up to 73.1% instance-level recovery and reduces FID and KL scores beyond two state-of-the-art baselines. Strikingly, on the challenging task of text unlearning, it improves BLEU scores by nearly 100% over black-box baselines while preserving distributional fidelity, demonstrating that RL can recover even high-dimensional and structured modalities. Furthermore, ReTrace demonstrates effectiveness across both convolutional (ResNet) and transformer-based models, with Distil-BERT as the largest architecture attacked to date. These results show that current unlearning methods remain vulnerable, highlighting the need for robust and provably private mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43658f00-50cb-45a8-98e6-ca57a2727d7bBuilds on15
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- Adaptive Machine UnlearningVarun Gupta, Christopher Jung, Seth Neel, Aaron Roth et al.NeurIPS 2021 · 262 citations
Related papers
- Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for PrivacyYaxin Xiao, Qingqing Ye, Li Hu, Huadi Zheng et al.ICCV 2025 · 6 citations
- Textual Unlearning Gives a False Sense of UnlearningJiacheng Du, Zhibo Wang, Jie Zhang, Xiaoyi Pang et al.ICML 2025
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu et al.ICLR 2026 · 15 citations
- Reconstruction Attacks on Machine Unlearning: Simple Models are VulnerableMartin Bertran Lopez, Shuai Tang, Michael Kearns, Jamie H. Morgenstern et al.NeurIPS 2024 · 39 citations
- Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine UnlearningHongsheng Hu, Shuo Wang, Tian Dong, Minhui XueS&P 2024 · 62 citations
