Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue
Jia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu, Zhen-Hua Ling, Kai-Wei Chang, Nanyun Peng
Abstract
Model editing is a technique that edits the large language models (LLMs) with updated knowledge to alleviate hallucinations without resource-intensive retraining. While current model editing methods can effectively modify a model's behavior within a specific area of interest, they often overlook the potential unintended side effects on the general abilities of LLMs such as reasoning, natural language inference, and question answering. In this paper, we raise concerns that model editing's improvements on factuality may come at the cost of a significant degradation of the model's general abilities. We systematically analyze the side effects by evaluating four popular editing methods on three LLMs across eight representative tasks. Our extensive empirical experiments show that it is challenging for current editing methods to simultaneously improve factuality of LLMs and maintain their general abilities. Our analysis reveals that the side effects are caused by model editing altering the original model weights excessively, leading to overfitting to the edited facts. To mitigate this, a method named RECT is proposed to regularize the edit update weights by imposing constraints on their complexity based on the RElative Change in weighT. Evaluation results show that RECT can significantly mitigate the side effects of editing while still maintaining over 94% editing performance 1 . * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39aef3d5-097e-48e6-a8eb-fa278ba53b46Cited by top-tier papers54
- The Mirage of Model Editing: Revisiting Evaluation in the WildWanli Yang, Fei Sun, Jiajun Tan, Xinyu Ma et al.ACL 2025 · 19 citations
- Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of RaceLihao Sun, Chengzhi Mao, Valentin Hofmann, Xuechunzi BaiACL 2025 · 13 citations
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation EditingYisong Xiao, Aishan Liu, Siyuan Liang, Zonghao Ying et al.NeurIPS 2025 · 12 citations
- Fine-tuning Done Right in Model EditingWanli Yang, Rui Tang, Hongyu Zang, Du Su et al.ICLR 2026 · 9 citations
- Rethinking Residual Distribution in Locate-then-Edit Model EditingXiaopeng Li, Shangwen Wang, Shasha Li, Shezheng Song et al.NeurIPS 2025 · 9 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 536 citations
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen et al.EMNLP 2023 · 449 citations
Related papers
- Should We Really Edit Language Models? On the Evaluation of Edited Language ModelsQi Li, Xiang Liu, Zhenheng Tang, Peijie Dong et al.NeurIPS 2024 · 25 citations
- Perturbation-Restrained Sequential Model EditingJun-Yu Ma, Hong Wang, Hao-Xiang Xu, Zhen-Hua Ling et al.ICLR 2025
- Navigating the Dual Facets: A Comprehensive Evaluation of Sequential Memory Editing in Large Language ModelsZihao Lin, Mohammad Beigi, Hongxuan Li, Yufan Zhou et al.ACL 2024 · 1 citation
- Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-FaithfulnessBaolong Bi, Shenghua Liu, Yiwei Wang, Lingrui Mei et al.ICLR 2025
- Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical EvidenceWanying Ren, Xin Song, Futing Wang, Guoxiu He et al.ICML 2026
