MULFE: A Multi-Level Benchmark for Free Text Model Editing
Chenhao Wang, Pengfei Cao, Zhuoran Jin, Yubo Chen, Daojian Zeng, Kang Liu, Jun Zhao
Abstract
Adjusting the outdated behaviors of large langugae models (LLMs) after deployment remains a significant challenge. It motivates the model editing research, which is however mainly explored in a restricted task form with triplebased edit requests. Recent works have initiated a transition to a more practical and unified editing task that takes free-form text as edit requests. However, there are gaps in nuanced benchmark designs and re-evaluation of existing methods. To bridge the gaps, we introduce a multi-level benchmark for free text model editing (MULFE). The benchmark categorizes probe queries into three levels of generalization, ranging from basic literal memory to deeper understanding and reasoning. Based on the benchmark, we conduct extensive experiments across various base models, edit sizes, and editing methods, including adaptations of mainstream locate-and-edit and hypernetwork methods. The results highlight the inconsistent behaviors of edited models on different generalization levels. Higherlevel generalization remains a significant challenge. Based on the findings, we propose SIDE, a simple yet effective method based on in-context distillation to enhance the generalization performance. The benchmark dataset and evaluation scripts are publicly available at http://github.com/wchrepo/mulfe .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73e7b191-c6d5-4a13-8c27-8611f1ad45faCited by top-tier papers3
- Revealing the Deceptiveness of Knowledge Editing: A Mechanistic Analysis of Superficial EditingJiakuan Xie, Pengfei Cao, Yubo Chen, Kang Liu et al.ACL 2025 · 2 citations
- Knowledge Localization: Mission Not Accomplished? Enter Query Localization!Yuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu et al.ICLR 2025
- M2Edit: Locate and Edit Multi-Granularity Knowledge in Multimodal Large Language ModelYang Zhou, Pengfei Cao, Yubo Chen, Qingbin Liu et al.EMNLP 2025
Builds on13
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni et al.ICLR 2024 · 462 citations
- Editable Neural NetworksAnton Sinitsin, Vsevolod Plokhotnyuk, Dmitry V. Pyrkin, Sergei Popov et al.ICLR 2020 · 210 citations
- Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMsOded Ovadia, Menachem Brief, Moshik Mishaeli, Oren ElishaEMNLP 2024 · 89 citations
- Editing Large Language Models: Problems, Methods, and OpportunitiesYunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng et al.EMNLP 2023 · 83 citations
- Massive Editing for Large Language Models via Meta LearningChenmien Tan, Ge Zhang, Jie FuICLR 2024 · 68 citations
Related papers
- Editing Across Languages: A Survey of Multilingual Knowledge EditingNadir Durrani, Basel Mousi, Fahim DalviEMNLP 2025
- Uncovering Overfitting in Large Language Model EditingMengqi Zhang, Xiaotian Ye, Qiang Liu, Shu Wu et al.ICLR 2025
- ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge EditingYaohui Ma, Xiaopeng Hong, Shizhou Zhang, Huiyun Li et al.AAAI 2025 · 2 citations
- Locate-then-edit for Multi-hop Factual Recall under Knowledge EditingZhuoran Zhang, Yongxiang Li, Zijian Kan, Keyuan Cheng et al.ICML 2025
- MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual KnowledgeYuntao Du, Kailin Jiang, Zhi Gao, Chenrui Shi et al.ICLR 2025
