MULFE: A Multi-Level Benchmark for Free Text Model Editing
Chenhao Wang, Pengfei Cao, Zhuoran Jin, Yubo Chen, Daojian Zeng, Kang Liu, Jun Zhao
摘要
Adjusting the outdated behaviors of large langugae models (LLMs) after deployment remains a significant challenge. It motivates the model editing research, which is however mainly explored in a restricted task form with triplebased edit requests. Recent works have initiated a transition to a more practical and unified editing task that takes free-form text as edit requests. However, there are gaps in nuanced benchmark designs and re-evaluation of existing methods. To bridge the gaps, we introduce a multi-level benchmark for free text model editing (MULFE). The benchmark categorizes probe queries into three levels of generalization, ranging from basic literal memory to deeper understanding and reasoning. Based on the benchmark, we conduct extensive experiments across various base models, edit sizes, and editing methods, including adaptations of mainstream locate-and-edit and hypernetwork methods. The results highlight the inconsistent behaviors of edited models on different generalization levels. Higherlevel generalization remains a significant challenge. Based on the findings, we propose SIDE, a simple yet effective method based on in-context distillation to enhance the generalization performance. The benchmark dataset and evaluation scripts are publicly available at http://github.com/wchrepo/mulfe .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Revealing the Deceptiveness of Knowledge Editing: A Mechanistic Analysis of Superficial EditingJiakuan Xie, Pengfei Cao, Yubo Chen, Kang Liu 等ACL 2025 · 被引用 2 次
- Knowledge Localization: Mission Not Accomplished? Enter Query Localization!Yuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu 等ICLR 2025
- M2Edit: Locate and Edit Multi-Granularity Knowledge in Multimodal Large Language ModelYang Zhou, Pengfei Cao, Yubo Chen, Qingbin Liu 等EMNLP 2025
它引用的顶会 Paper13
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni 等ICLR 2024 · 被引用 462 次
- Editable Neural NetworksAnton Sinitsin, Vsevolod Plokhotnyuk, Dmitry V. Pyrkin, Sergei Popov 等ICLR 2020 · 被引用 210 次
- Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMsOded Ovadia, Menachem Brief, Moshik Mishaeli, Oren ElishaEMNLP 2024 · 被引用 89 次
- Editing Large Language Models: Problems, Methods, and OpportunitiesYunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng 等EMNLP 2023 · 被引用 83 次
- Massive Editing for Large Language Models via Meta LearningChenmien Tan, Ge Zhang, Jie FuICLR 2024 · 被引用 68 次
相关 Paper
- Editing Across Languages: A Survey of Multilingual Knowledge EditingNadir Durrani, Basel Mousi, Fahim DalviEMNLP 2025
- Uncovering Overfitting in Large Language Model EditingMengqi Zhang, Xiaotian Ye, Qiang Liu, Shu Wu 等ICLR 2025
- ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge EditingYaohui Ma, Xiaopeng Hong, Shizhou Zhang, Huiyun Li 等AAAI 2025 · 被引用 2 次
- Locate-then-edit for Multi-hop Factual Recall under Knowledge EditingZhuoran Zhang, Yongxiang Li, Zijian Kan, Keyuan Cheng 等ICML 2025
- MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual KnowledgeYuntao Du, Kailin Jiang, Zhi Gao, Chenrui Shi 等ICLR 2025
