Identifying Knowledge Editing Types in Large Language Models
Xiaopeng Li, Shasha Li, Shangwen Wang, Shezheng Song, Bin Ji, Huijun Liu, Jun Ma, Jie Yu
Abstract
Warning: This paper contains examples of toxic text.
Knowledge editing has emerged as an efficient technique for updating the knowledge of large language models (LLMs), attracting increasing attention in recent years. However, there is a lack of effective measures to prevent the malicious misuse of this technique, which could lead to harmful edits in LLMs. These malicious modifications could cause LLMs to generate toxic content, misleading users into inappropriate actions. In front of this risk, we introduce a new task, Knowledge Editing Type Identification (KETI), aimed at identifying different types of edits in LLMs, thereby providing timely alerts to users when encountering illicit edits. As part of this task, we propose KETIBench, which includes five types of harmful edits covering the most popular toxic types, as well as one benign factual edit. We develop five classical classification models and three BERT-based models as baseline identifiers for both open-source and closed-source LLMs. Our experimental results, across 92 trials involving four models and three knowledge editing methods, demonstrate that all eight baseline identifiers achieve decent identification performance, highlighting the feasibility of identifying malicious edits in LLMs. Additional analyses reveal that the performance of the identifiers is independent of the reliability of the knowledge editing methods and exhibits cross-domain generalization, enabling the identification of edits from unknown sources. All data and code are available in https://github.com/xpq-tech/KETI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 016b80c9-05a4-47db-bfb2-dd735c76bc15Cited by top-tier papers3
- Tracing and Reversing Edits in LLMsPaul Youssef, Zhixue Zhao, Christin Seifert, Jörg SchlöttererICLR 2026 · 7 citations
- Can Fine-Tuning Erase Edits? On the Fragile Coexistence of Knowledge Editing and Fine-tuningYinjie Cheng, Paul Youssef, Christin Seifert, Jörg Schlötterer et al.KDD 2026 · 2 citations
- Can Knowledge Editing Really Correct Hallucinations?Baixiang Huang, Canyu Chen, Xiongxiao Xu, Ali Payani et al.ICLR 2025
Builds on17
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
- Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsTom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim et al.NeurIPS 2023 · 349 citations
- PMET: Precise Model Editing in a TransformerXiaopeng Li, Shasha Li, Shezheng Song, Jing Yang et al.AAAI 2024 · 208 citations
- WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language ModelsPeng Wang, Zexi Li, Ningyu Zhang, Ziwen Xu et al.NeurIPS 2024 · 125 citations
Related papers
- Detoxifying Large Language Models via Knowledge EditingMengru Wang, Ningyu Zhang, Ziwen Xu, Zekun Xi et al.ACL 2024
- Unveiling the Pitfalls of Knowledge Editing for Large Language ModelsZhoubo Li, Ningyu Zhang, Yunzhi Yao, Mengru Wang et al.ICLR 2024 · 47 citations
- MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual KnowledgeYuntao Du, Kailin Jiang, Zhi Gao, Chenrui Shi et al.ICLR 2025
- Can Editing LLMs Inject Harm?Canyu Chen, Baixiang Huang, Zekun Li, Zhaorun Chen et al.AAAI 2026 · 26 citations
- Assessing and Post-Processing Black Box Large Language Models for Knowledge EditingXiaoshuai Song, Zhengyang Wang, Keqing He, Guanting Dong et al.WWW 2025 · 3 citations
