Revealing the Deceptiveness of Knowledge Editing: A Mechanistic Analysis of Superficial Editing
Jiakuan Xie, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
Abstract
Knowledge editing, which aims to update the knowledge encoded in language models, can be deceptive. Despite the fact that many existing knowledge editing algorithms achieve near-perfect performance on conventional metrics, the models edited by them are still prone to generating original knowledge. This paper introduces the concept of "superficial editing" to describe this phenomenon. Our comprehensive evaluation reveals that this issue presents a significant challenge to existing algorithms. Through systematic investigation, we identify and validate two key factors contributing to this issue: (1) the residual stream at the last subject position in earlier layers and (2) specific attention modules in later layers. Notably, certain attention heads in later layers, along with specific left singular vectors in their output matrices, encapsulate the original knowledge and exhibit a causal relationship with superficial editing. Furthermore, we extend our analysis to the task of superficial unlearning, where we observe consistent patterns in the behavior of specific attention heads and their corresponding left singular vectors, thereby demonstrating the robustness and broader applicability of our methodology and conclusions. Our code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical EvidenceWanying Ren, Xin Song, Futing Wang, Guoxiu He et al.ICML 2026
- Cracking Factual Knowledge: A Comprehensive Analysis of Degenerate Knowledge Neurons in Large Language ModelsYuheng Chen, Pengfei Cao, Yubo Chen, Yining Wang et al.ACL 2025
- M2Edit: Locate and Edit Multi-Granularity Knowledge in Multimodal Large Language ModelYang Zhou, Pengfei Cao, Yubo Chen, Qingbin Liu et al.EMNLP 2025
Builds on21
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsTom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim et al.NeurIPS 2023 · 349 citations
- PMET: Precise Model Editing in a TransformerXiaopeng Li, Shasha Li, Shezheng Song, Jing Yang et al.AAAI 2024 · 208 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
- Overthinking the Truth: Understanding how Language Models Process False DemonstrationsDanny Halawi, Jean-Stanislas Denain, Jacob SteinhardtICLR 2024 · 83 citations
Related papers
- Spectral Characterization and Mitigation of Sequential Knowledge Editing CollapseChi Zhang, Mengqi Zhang, Xiaotian Ye, Runxi Cheng et al.ACL 2026 · 2 citations
- AdaEdit: Advancing Continuous Knowledge Editing For Large Language ModelsQi Li, Xiaowen ChuACL 2025
- Should We Really Edit Language Models? On the Evaluation of Edited Language ModelsQi Li, Xiang Liu, Zhenheng Tang, Peijie Dong et al.NeurIPS 2024 · 25 citations
- Revealing and Mitigating Over-Attention in Knowledge EditingPinzheng Wang, Zecheng Tang, Keyan Zhou, Juntao Li et al.ICLR 2025
- FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of KnowledgeNakyeong Yang, Minsung Kim, Seunghyun Yoon, Joongbo Shin et al.EMNLP 2025
