REMIND: Memorization and Unlearning in LLMs Through the Lens of Input Loss Landscapes
Liran Cohen, Yaniv Nemcovsky, Avi Mendelson
摘要
Understanding how large language models (LLMs) store, retain, and remove knowledge is critical for interpretability, reliability, and privacy compliance. We reveal a key phenomenon: machine unlearning imprints distinct geometric signatures in the model's input loss landscape (ILL), with unlearned examples forming flat, low-curvature plateaus contrasting the sharp, high-curvature basins of retained or unseen examples, even when pointwise losses overlap, exposing residual memorization through inputoutput behavior alone. Building on this, we introduce REMIND (Residual Memorization in Neighborhood Dynamics), a framework that diagnoses memorization states (retained, forgotten, holdout) by probing local ILL curvature over semantically coherent neighborhoods, using only loss queries and a novel embeddingproximity perturbation method for generating controlled, interpretable variants. REMIND achieves 82% multi-class ROC-AUC in aggregate evaluations, outperforming baselines like ROUGE-L and MIN-K%++, with roughly 2x higher AUC at 1% FPR, and remains robust on paraphrased inputs. This neighborhood-level geometric analysis provides a practical, interpretable lens on LLM knowledge retention and unlearning, detecting subtle residual signals missed by pointwise or aggregated metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang 等ICLR 2024 · 被引用 365 次
- Simplicity Prevails: Rethinking Negative Preference Optimization for LLM UnlearningChongyu Fan, Jiancheng Liu, Licong Lin, Jinghan Jia 等NeurIPS 2025 · 被引用 182 次
- Erasing Conceptual Knowledge from Language ModelsRohit Gandikota, Sheridan Feucht, Samuel Marks, David BauNeurIPS 2025 · 被引用 35 次
- Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt CalibrationWenjie Fu, Huandong Wang, Chen Gao, Guanghua Liu 等NeurIPS 2024 · 被引用 28 次
相关 Paper
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu 等ICLR 2026 · 被引用 15 次
- Towards Effective Evaluations and Comparisons for LLM Unlearning MethodsQizhou Wang, Bo Han, Puning Yang, Jianing Zhu 等ICLR 2025
- Model Unlearning via Sparse Autoencoder Subspace Guided ProjectionsXu Wang, Zihao Li, Benyou Wang, Yan Hu 等EMNLP 2025 · 被引用 9 次
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMsXiaoyu Xu, Xiang Yue, Yang Liu, Qingqing Ye 等ICML 2026 · 被引用 36 次
- Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence AwarenessRongzhe Wei, Peizhi Niu, Hans Hao-Hsun Hsu, Ruihan Wu 等NeurIPS 2025 · 被引用 20 次
