Representation Unlearning: Forgetting through Information Compression
Antonio Almudévar, Alfonso Ortega
Abstract
Machine unlearning seeks to remove the influence of specific training data from a model, a need driven by privacy regulations and robustness concerns. Existing approaches typically modify model parameters, but such updates can be unstable, computationally costly, and limited by local approximations. We introduce Representation Unlearning, a framework that performs unlearning directly in the model’s representation space. Instead of modifying model parameters, we learn a transformation over representations that imposes an information bottleneck: maximizing mutual information with retained data while suppressing information about data to be forgotten. We derive variational surrogates that make this objective tractable and show how they can be instantiated in two practical regimes: when both retain and forget data are available, and in a zero-shot setting where only forget data can be accessed. Experiments across several benchmarks demonstrate that Representation Unlearning achieves more reliable forgetting, better utility retention, and greater computational efficiency than parameter-centric baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfc8729a-99df-4334-8e5c-2d14d657199dBuilds on21
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Exploiting Unintended Feature Leakage in Collaborative LearningLuca Melis, Congzheng Song, Emiliano De Cristofaro, Vitaly ShmatikovS&P 2019 · 1,736 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 516 citations
Related papers
- Unlearning-Aware MinimizationHoki Kim, Keonwoo Kim, Sungwon Chae, Sangwon YoonNeurIPS 2025 · 7 citations
- ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language ModelsYujie Lin, Chengyi Yang, Zhishang Xiang, YIPING SONG et al.ICML 2026
- Label-Agnostic Forgetting: A Supervision-Free Unlearning in Deep ModelsShaofei Shen, Chenhao Zhang, Yawen Zhao, Alina Bialkowski et al.ICLR 2024 · 20 citations
- The Utility and Complexity of In- and Out-of-Distribution Machine UnlearningYoussef Allouah, Joshua Kazdan, Rachid Guerraoui, Sanmi KoyejoICLR 2025
- MUter: Machine Unlearning on Adversarially Trained ModelsJunxu Liu, Mingsheng Xue, Jian Lou, Xiaoyu Zhang et al.ICCV 2023 · 36 citations
