Tracing and Reversing Edits in LLMs
Paul Youssef, Zhixue Zhao, Christin Seifert, Jörg Schlötterer
Abstract
Knowledge editing methods (KEs) are a cost-effective way to update the factual content of large language models (LLMs), but they pose a dual-use risk. While KEs are beneficial for updating outdated or incorrect information, they can be exploited maliciously to implant misinformation or bias. In order to defend against these types of malicious manipulation, we need robust techniques that can reliably detect, interpret, and mitigate malicious edits. To that end, we introduce the tasks of tracing and reversing edits. We propose a novel method to infer the edited object entity, solely based on the modified weights, without access to the editing prompt or any other semantically similar prompts, with up to 99% accuracy. Further, we propose an effective and training-free method for reversing edits. Our method reverses up to 94% of the edits, and helps regain the original model's output distribution without access to any information about the edit. This method can further be repurposed to distinguish between edited and unedited weights. Our findings highlight the feasibility of tracing and reversing edits based on the edited weights, opening a new research direction for safeguarding LLMs against adversarial manipulations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28a6546a-97d0-45c6-a42b-1592408e35ccCited by top-tier papers2
- Can Fine-Tuning Erase Edits? On the Fragile Coexistence of Knowledge Editing and Fine-tuningYinjie Cheng, Paul Youssef, Christin Seifert, Jörg Schlötterer et al.KDD 2026 · 2 citations
- Reverse-Engineering Model Editing on Language ModelsZhiyu Sun, Minrui Luo, Yu Wang, Zhili Chen et al.ICML 2026 · 1 citation
Builds on13
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning et al.ICML 2022 · 520 citations
- WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language ModelsPeng Wang, Zexi Li, Ningyu Zhang, Ziwen Xu et al.NeurIPS 2024 · 125 citations
- BadEdit: Backdooring Large Language Models by Model EditingYanzhou Li, Tianlin Li, Kangjie Chen, Jian Zhang et al.ICLR 2024 · 116 citations
Related papers
- Can Editing LLMs Inject Harm?Canyu Chen, Baixiang Huang, Zekun Li, Zhaorun Chen et al.AAAI 2026 · 26 citations
- Identifying Knowledge Editing Types in Large Language ModelsXiaopeng Li, Shasha Li, Shangwen Wang, Shezheng Song et al.KDD 2025
- AdaEdit: Advancing Continuous Knowledge Editing For Large Language ModelsQi Li, Xiaowen ChuACL 2025
- Large Scale Knowledge WashingYu Wang, Ruihan Wu, Zexue He, Xiusi Chen et al.ICLR 2025
- Unveiling the Pitfalls of Knowledge Editing for Large Language ModelsZhoubo Li, Ningyu Zhang, Yunzhi Yao, Mengru Wang et al.ICLR 2024 · 47 citations
