Tracing and Reversing Edits in LLMs
Paul Youssef, Zhixue Zhao, Christin Seifert, Jörg Schlötterer
摘要
Knowledge editing methods (KEs) are a cost-effective way to update the factual content of large language models (LLMs), but they pose a dual-use risk. While KEs are beneficial for updating outdated or incorrect information, they can be exploited maliciously to implant misinformation or bias. In order to defend against these types of malicious manipulation, we need robust techniques that can reliably detect, interpret, and mitigate malicious edits. To that end, we introduce the tasks of tracing and reversing edits. We propose a novel method to infer the edited object entity, solely based on the modified weights, without access to the editing prompt or any other semantically similar prompts, with up to 99% accuracy. Further, we propose an effective and training-free method for reversing edits. Our method reverses up to 94% of the edits, and helps regain the original model's output distribution without access to any information about the edit. This method can further be repurposed to distinguish between edited and unedited weights. Our findings highlight the feasibility of tracing and reversing edits based on the edited weights, opening a new research direction for safeguarding LLMs against adversarial manipulations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Can Fine-Tuning Erase Edits? On the Fragile Coexistence of Knowledge Editing and Fine-tuningYinjie Cheng, Paul Youssef, Christin Seifert, Jörg Schlötterer 等KDD 2026 · 被引用 2 次
- Reverse-Engineering Model Editing on Language ModelsZhiyu Sun, Minrui Luo, Yu Wang, Zhili Chen 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper13
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn 等ICLR 2022 · 被引用 527 次
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning 等ICML 2022 · 被引用 520 次
- WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language ModelsPeng Wang, Zexi Li, Ningyu Zhang, Ziwen Xu 等NeurIPS 2024 · 被引用 125 次
- BadEdit: Backdooring Large Language Models by Model EditingYanzhou Li, Tianlin Li, Kangjie Chen, Jian Zhang 等ICLR 2024 · 被引用 116 次
相关 Paper
- Can Editing LLMs Inject Harm?Canyu Chen, Baixiang Huang, Zekun Li, Zhaorun Chen 等AAAI 2026 · 被引用 26 次
- Identifying Knowledge Editing Types in Large Language ModelsXiaopeng Li, Shasha Li, Shangwen Wang, Shezheng Song 等KDD 2025
- AdaEdit: Advancing Continuous Knowledge Editing For Large Language ModelsQi Li, Xiaowen ChuACL 2025
- Large Scale Knowledge WashingYu Wang, Ruihan Wu, Zexue He, Xiusi Chen 等ICLR 2025
- Unveiling the Pitfalls of Knowledge Editing for Large Language ModelsZhoubo Li, Ningyu Zhang, Yunzhi Yao, Mengru Wang 等ICLR 2024 · 被引用 47 次
