REVIVING YOUR MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
Aly M. Kassem, Zhuan Shi, Negar Rostamzadeh, Golnoosh Farnadi
Abstract
LLMs are frequently fine-tuned or unlearned to adapt to new tasks or eliminate undesirable behaviors. While existing evaluation methods assess performance after such interventions, there remains no general approach for detecting unintended side effects-such as unlearning biology content degrading performance on chemistry tasks, particularly when these effects are unpredictable or emergent. To address this issue, we introduce MNEME, Model diffiNg for Evaluating Mechanistic Effects, a framework for identifying these side effects using sparse model diffing. MNEME compares base and fine-tuned models on out-of-distribution (OOD) data (e.g., The Pile, LMSYS-Chat-1M), without access to fine-tuning data, to isolate behavioral shifts. Applied to five LLMs across three scenarios, WMDP knowledge unlearning, emergent misalignment, and benign finetuning, MNEME achieves up to 95% accuracy in predicting side effects, aligning with known benchmarks and requiring no custom heuristics. Our results demonstrate that sparse probing and diffing offer a scalable and automated lens into fine-tuning-induced model changes, providing practical tools for understanding and managing LLM behavior. 1 * MNEME refers to Mnēmosynē, the Greek Titan goddess of memory, whose name derives from the Greek word mnēmē ("memory").
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d22f4e44-83c2-4c27-8dea-270d4a12fe7aCited by top-tier papers3
- Steering Out-of-Distribution Generalization with Concept Ablation Fine-TuningHelena Casademunt, Caden Juang, Adam Karvonen, Samuel Marks et al.ICML 2026 · 32 citations
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 18 citations
- Med-SegLens: Latent-Level Model Diffing for Interpretable Medical Image SegmentationSalma Ahmed, Emad Mohammed, Azam BidgoliICML 2026 · 1 citation
Builds on8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingNikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov et al.ICLR 2024 · 113 citations
- Mechanistically analyzing the effects of fine-tuning on procedurally defined tasksSamyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P. Dick et al.ICLR 2024 · 108 citations
- Preserving Privacy Through Dememorization: An Unlearning Technique For Mitigating Memorization Risks In Language ModelsAly M. Kassem, Omar Mahmoud, Sherif SaadEMNLP 2023 · 10 citations
- Model Editing Harms General Abilities of Large Language Models: Regularization to the RescueJia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu et al.EMNLP 2024 · 8 citations
Related papers
- Narrow Finetuning Leaves Clearly Readable Traces in Activation DifferencesJulian Minder, Clément Dumas, Stewart Slocum, Helena Casademunt et al.ICLR 2026 · 29 citations
- Learning to Interpret Weight Differences in Language ModelsAvichal Goel, Yoon Kim, Nir N Shavit, Tony T. WangICLR 2026 · 10 citations
- Mitigating Memorization in Language ModelsMansi Sakarvadia, Aswathy Ajith, Arham Mushtaq Khan, Nathaniel C. Hudson et al.ICLR 2025
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMsZiqian Zhong, Aditi RaghunathanICLR 2026 · 7 citations
- REMIND: Memorization and Unlearning in LLMs Through the Lens of Input Loss LandscapesLiran Cohen, Yaniv Nemcovsky, Avi MendelsonACL 2026
