Massive Editing for Large Language Models via Meta Learning
Chenmien Tan, Ge Zhang, Jie Fu
Abstract
While large language models (LLMs) have enabled learning knowledge from the pre-training corpora, the acquired knowledge may be fundamentally incorrect or outdated over time, which necessitates rectifying the knowledge of the language model (LM) after the training. A promising approach involves employing a hyper-network to generate parameter shift, whereas existing hyper-networks suffer from inferior scalability in synchronous editing operation amount (Hase et al., 2023b; Huang et al., 2023) . For instance, Mitchell et al. ( 2022 ) mimic gradient accumulation to sum the parameter shifts together, which lacks statistical significance and is prone to cancellation effect. To mitigate the problem, we propose the MAssive Language Model Editing Network (MALMEN), which formulates the parameter shift aggregation as the least square problem, subsequently updating the LM parameters using the normal equation. To accommodate editing multiple facts simultaneously with limited memory budgets, we separate the computation on the hyper-network and LM, enabling arbitrary batch size on both neural networks. Our method is evaluated by editing up to thousands of facts on LMs with different architectures, i.e., BERT-base, GPT-2, T5-XL (2.8B), and GPT-J (6B), across various knowledge-intensive NLP tasks, i.e., closed book fact-checking and question answering. Remarkably, MALMEN is capable of editing hundreds of times more facts than MEND (Mitchell et al., 2022) with the identical hyper-network architecture and outperforms editor specifically designed for GPT, i.e., MEMIT (Meng et al.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c506c58b-fd19-4c26-965a-81bfefc1f3bdCited by top-tier papers57
- Self-Adapting Language ModelsAdam Zweiger, Jyothish Pari, Han Guo, Yoon Kim et al.NeurIPS 2025 · 78 citations
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li et al.ACL 2025 · 33 citations
- The Mirage of Model Editing: Revisiting Evaluation in the WildWanli Yang, Fei Sun, Jiajun Tan, Xinyu Ma et al.ACL 2025 · 19 citations
- Reasons and Solutions for the Decline in Model Performance after EditingXiusheng Huang, Jiaxiang Liu, Yequan Wang, Kang LiuNeurIPS 2024 · 13 citations
- SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding AlteringXiaopeng Li, Shasha Li, Shezheng Song, Huijun Liu et al.AAAI 2025 · 11 citations
Builds on14
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
Related papers
- Editing Factual Knowledge in Language ModelsNicola De Cao, Wilker Aziz, Ivan TitovEMNLP 2021 · 20 citations
- EAMET: Robust Massive Model Editing via Embedding Alignment OptimizationYanbo Dai, Zhenlan Ji, Zongjie Li, Shuai WangICLR 2026
- Knowledge Graph Enhanced Large Language Model EditingMengqi Zhang, Xiaotian Ye, Qiang Liu, Pengjie Ren et al.EMNLP 2024 · 5 citations
- Editing Across Languages: A Survey of Multilingual Knowledge EditingNadir Durrani, Basel Mousi, Fahim DalviEMNLP 2025
- Scaling Knowledge Editing in LLMs to 100, 000 Facts with Neural KV DatabaseWeizhi Fei, Hao Shi, Jing Xu, Jingchen Peng et al.ICLR 2026 · 2 citations
