REVIVING YOUR MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
Aly M. Kassem, Zhuan Shi, Negar Rostamzadeh, Golnoosh Farnadi
摘要
LLMs are frequently fine-tuned or unlearned to adapt to new tasks or eliminate undesirable behaviors. While existing evaluation methods assess performance after such interventions, there remains no general approach for detecting unintended side effects-such as unlearning biology content degrading performance on chemistry tasks, particularly when these effects are unpredictable or emergent. To address this issue, we introduce MNEME, Model diffiNg for Evaluating Mechanistic Effects, a framework for identifying these side effects using sparse model diffing. MNEME compares base and fine-tuned models on out-of-distribution (OOD) data (e.g., The Pile, LMSYS-Chat-1M), without access to fine-tuning data, to isolate behavioral shifts. Applied to five LLMs across three scenarios, WMDP knowledge unlearning, emergent misalignment, and benign finetuning, MNEME achieves up to 95% accuracy in predicting side effects, aligning with known benchmarks and requiring no custom heuristics. Our results demonstrate that sparse probing and diffing offer a scalable and automated lens into fine-tuning-induced model changes, providing practical tools for understanding and managing LLM behavior. 1 * MNEME refers to Mnēmosynē, the Greek Titan goddess of memory, whose name derives from the Greek word mnēmē ("memory").
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Steering Out-of-Distribution Generalization with Concept Ablation Fine-TuningHelena Casademunt, Caden Juang, Adam Karvonen, Samuel Marks 等ICML 2026 · 被引用 32 次
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 被引用 18 次
- Med-SegLens: Latent-Level Model Diffing for Interpretable Medical Image SegmentationSalma Ahmed, Emad Mohammed, Azam BidgoliICML 2026 · 被引用 1 次
它引用的顶会 Paper8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingNikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov 等ICLR 2024 · 被引用 113 次
- Mechanistically analyzing the effects of fine-tuning on procedurally defined tasksSamyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P. Dick 等ICLR 2024 · 被引用 108 次
- Preserving Privacy Through Dememorization: An Unlearning Technique For Mitigating Memorization Risks In Language ModelsAly M. Kassem, Omar Mahmoud, Sherif SaadEMNLP 2023 · 被引用 10 次
- Model Editing Harms General Abilities of Large Language Models: Regularization to the RescueJia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu 等EMNLP 2024 · 被引用 8 次
相关 Paper
- Narrow Finetuning Leaves Clearly Readable Traces in Activation DifferencesJulian Minder, Clément Dumas, Stewart Slocum, Helena Casademunt 等ICLR 2026 · 被引用 29 次
- Learning to Interpret Weight Differences in Language ModelsAvichal Goel, Yoon Kim, Nir N Shavit, Tony T. WangICLR 2026 · 被引用 10 次
- Mitigating Memorization in Language ModelsMansi Sakarvadia, Aswathy Ajith, Arham Mushtaq Khan, Nathaniel C. Hudson 等ICLR 2025
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMsZiqian Zhong, Aditi RaghunathanICLR 2026 · 被引用 7 次
- REMIND: Memorization and Unlearning in LLMs Through the Lens of Input Loss LandscapesLiran Cohen, Yaniv Nemcovsky, Avi MendelsonACL 2026
