Uncovering Overfitting in Large Language Model Editing
Mengqi Zhang, Xiaotian Ye, Qiang Liu, Shu Wu, Pengjie Ren, Zhumin Chen
Abstract
Knowledge editing has been proposed as an effective method for updating and correcting the internal knowledge of Large Language Models (LLMs). However, existing editing methods often struggle with complex tasks, such as multi-hop reasoning. In this paper, we identify and investigate the phenomenon of Editing Overfit, where edited models assign disproportionately high probabilities to the edit target, hindering the generalization of new knowledge in complex scenarios. We attribute this issue to the current editing paradigm, which places excessive emphasis on the direct correspondence between the input prompt and the edit target for each edit sample. To further explore this issue, we introduce a new benchmark, EVOKE (EValuation of Editing Overfit in Knowledge Editing), along with fine-grained evaluation metrics. Through comprehensive experiments and analysis, we demonstrate that Editing Overfit is prevalent in current editing methods and that common overfitting mitigation strategies are ineffective in knowledge editing. To overcome this, inspired by LLMs' knowledge recall mechanisms, we propose a new plug-and-play strategy called Learn the Inference (LTI), which introduce a Multi-stage Inference Constraint module to guide the edited models in recalling new knowledge similarly to how unedited LLMs leverage knowledge through in-context learning. Extensive experimental results across a wide range of tasks validate the effectiveness of LTI in mitigating Editing Overfit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b467e726-c566-4acf-9137-4a8cbd2a5a78Cited by top-tier papers7
- Tracing and Reversing Edits in LLMsPaul Youssef, Zhixue Zhao, Christin Seifert, Jörg SchlöttererICLR 2026 · 7 citations
- Disentangling Knowledge Representations for Large Language Model EditingMengqi Zhang, Zisheng Zhou, Xiaotian Ye, Qiang Liu et al.ICLR 2026 · 6 citations
- SAKE: Steering Activations for Knowledge EditingMarco Scialanga, Thibault Laugel, Vincent Grari, Marcin DetynieckiACL 2025 · 6 citations
- LLM Unlearning Should Be Form-IndependentXiaotian Ye, Mengqi Zhang, Shu WuS&P 2026 · 3 citations
- Mitigating Heterogeneous Token Overfitting in LLM Knowledge EditingTianci Liu, Ruirui Li, Zihan Dong, Hui Liu et al.ICML 2025
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
- Editing Large Language Models: Problems, Methods, and OpportunitiesYunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng et al.EMNLP 2023 · 83 citations
Related papers
- Locate-then-edit for Multi-hop Factual Recall under Knowledge EditingZhuoran Zhang, Yongxiang Li, Zijian Kan, Keyuan Cheng et al.ICML 2025
- History Matters: Temporal Knowledge Editing in Large Language ModelXunjian Yin, Jin Jiang, Liming Yang, Xiaojun WanAAAI 2024 · 18 citations
- CaKE: Circuit-aware Editing Enables Generalizable Knowledge LearnersYunzhi Yao, Jizhan Fang, Jia-Chen Gu, Ningyu Zhang et al.EMNLP 2025 · 1 citation
- Learning to Edit: Aligning LLMs with Knowledge EditingYuxin Jiang, Yufei Wang, Chuhan Wu, Wanjun Zhong et al.ACL 2024 · 9 citations
- REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge EditingHaitian Zhong, Yuhuan Liu, Ziyang Xu, Guofan Liu et al.EMNLP 2025
