Forget What Matters, Keep the Rest: Selective Unlearning of Informative Tokens
Seunghee Koh, Sunghyun Baek, Youngdong Kim, Junmo Kim
Abstract
Unlearning in large language models (LLMs) has emerged as a promising safeguard against adversarial behaviors. When the forgetting loss is applied uniformly without considering tokenlevel semantic importance, model utility can be unnecessarily degraded. Recent studies have explored token-wise loss regularizers that prioritize informative tokens, but largely rely on ground-truth confidence or external linguistic parsers, which limits their ability to capture contextual information or the model's overall predictive state. Intuitively, function words like "the" primarily serve syntactic roles and are highly predictable with little ambiguity, but informative words admit multiple plausible alternatives with greater uncertainty. Based on this intuition, we propose Entropy-guided Token Weighting (ETW), a token-level unlearning regularizer that uses entropy of the predictive distribution as a proxy for token informativeness. We demonstrate that informative tokens tend to have higher entropy, whereas structural tokens tend to have lower entropy. This behavior enables ETW to achieve more effective unlearning while better preserving model utility than existing token-level approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55a6b71d-aa37-4923-9415-4a86293c29a1Builds on13
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang et al.ICLR 2024 · 365 citations
- Reasoning with Exploration: An Entropy PerspectiveDaixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai et al.AAAI 2026 · 216 citations
- Simplicity Prevails: Rethinking Negative Preference Optimization for LLM UnlearningChongyu Fan, Jiancheng Liu, Licong Lin, Jinghan Jia et al.NeurIPS 2025 · 182 citations
Related papers
- Confidence Regulation Neurons in Language ModelsAlessandro Stolfo, Ben Wu, Wes Gurnee, Yonatan Belinkov et al.NeurIPS 2024 · 68 citations
- Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM UnlearningNaixin Zhai, Pengyang Shao, Binbin Zheng, Yonghui Yang et al.ACL 2026 · 10 citations
- A Closer Look at Machine Unlearning for Large Language ModelsXiaojian Yuan, Tianyu Pang, Chao Du, Kejiang Chen et al.ICLR 2025
- WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language ModelsJinghan Jia, Jiancheng Liu, Yihua Zhang, Parikshit Ram et al.NeurIPS 2024 · 32 citations
- ALTER: Asymmetric LoRA for Token-Entropy-Guided Unlearning of LLMsXunlei Chen, Jinyu Guo, Yuang Li, Zhaokun Wang et al.AAAI 2026 · 2 citations
