Redirection for Erasing Memory (REM): Towards a universal unlearning method for corrupted data
Stefan Schoepf, Michael Mozer, Nicole Mitchell, Alexandra Brintrup, Georgios Kaissis, Peter Kairouz, Eleni Triantafillou
Abstract
Machine unlearning is studied for a multitude of tasks, but specialization of unlearning methods to particular tasks has made their systematic comparison challenging. To address this issue, we propose a conceptual space to characterize diverse corrupted data unlearning tasks in vision classifiers. This space is described by two dimensions, the discovery rate (the fraction of the corrupted data that are known at unlearning time) and the statistical regularity of the corrupted data (from random exemplars to shared concepts). Methods proposed previously have been targeted at portions of this space and-we show-fail predictably outside these regions. We propose a novel method, Redirection for Erasing Memory (REM), whose key feature is that corrupted data are redirected to dedicated neurons introduced at unlearning time and then discarded or deactivated to suppress the influence of corrupted data. REM performs strongly across the space of tasks, in contrast to prior SOTA methods that fail outside the regions for which they were designed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29337174-905b-4da6-b378-c48b5ccc52d2Cited by top-tier papers2
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMsKyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak et al.ICLR 2026 · 59 citations
- Variance-Reduced Unlearning using Forget Set GradientsMartin Van Waerebeke, Giovanni Neglia, Kevin Scaman, Marco Lorenzi et al.ICML 2026
Builds on16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 516 citations
- Towards Unbounded Machine UnlearningMeghdad Kurmanji, Peter Triantafillou, Jamie Hayes, Eleni TriantafillouNeurIPS 2023 · 363 citations
Related papers
- ERM-KTP: Knowledge-Level Machine Unlearning via Knowledge TransferShen Lin, Xiaoyu Zhang, Chenyang Chen, Xiaofeng Chen et al.CVPR 2023
- What makes unlearning hard and what to do about itKairan Zhao, Meghdad Kurmanji, George-Octavian Barbulescu, Eleni Triantafillou et al.NeurIPS 2024 · 115 citations
- Tackling Fake Forgetting through Uncertainty QuantificationYingdan Shi, Sijia Liu, Kaize Ding, Ren WangICML 2026 · 1 citation
- FUNU: Boosting Machine Unlearning Efficiency by Filtering Unnecessary UnlearningZitong Li, Qingqing Ye, Haibo HuWWW 2025 · 8 citations
- CoUn: Empowering Machine Unlearning via Contrastive LearningYasser H. Khalil, Mehdi Setayesh, Hongliang LiNeurIPS 2025 · 4 citations
