PMET: Precise Model Editing in a Transformer
Xiaopeng Li, Shasha Li, Shezheng Song, Jing Yang, Jun Ma, Jie Yu
摘要
Model editing techniques modify a minor proportion of knowledge in Large Language Models (LLMs) at a relatively low cost, which have demonstrated notable success. Existing methods assume Transformer Layer (TL) hidden states are values of key-value memories of the Feed-Forward Network (FFN). They usually optimize the TL hidden states to memorize target knowledge and use it to update the weights of the FFN in LLMs. However, the information flow of TL hidden states comes from three parts: Multi-Head Self-Attention (MHSA), FFN, and residual connections. Existing methods neglect the fact that the TL hidden states contains information not specifically required for FFN. Consequently, the performance of model editing decreases. To achieve more precise model editing, we analyze hidden states of MHSA and FFN, finding that MHSA encodes certain general knowledge extraction patterns. This implies that MHSA weights do not require updating when new knowledge is introduced. Based on above findings, we introduce PMET, which simultaneously optimizes Transformer Component (TC, namely MHSA and FFN) hidden states, while only using the optimized TC hidden states of FFN to precisely update FFN weights. Our experiments demonstrate that PMET exhibits state-of-the-art performance on both the counterfact and zsRE datasets. Our ablation experiments substantiate the effectiveness of our enhancements, further reinforcing the finding that the MHSA encodes certain general knowledge extraction patterns and indicating its storage of a small amount of factual knowledge. Our code is available at https://github.com/xpq-tech/PMET.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper74
- Teach AI How to Code: Using Large Language Models as Teachable Agents for Programming EducationHyoungwook Jin, Seonghee Lee, Hyungyu Shin, Juho KimCHI 2024 · 被引用 94 次
- Unveiling the Pitfalls of Knowledge Editing for Large Language ModelsZhoubo Li, Ningyu Zhang, Yunzhi Yao, Mengru Wang 等ICLR 2024 · 被引用 47 次
- Can Editing LLMs Inject Harm?Canyu Chen, Baixiang Huang, Zekun Li, Zhaorun Chen 等AAAI 2026 · 被引用 26 次
- The Mirage of Model Editing: Revisiting Evaluation in the WildWanli Yang, Fei Sun, Jiajun Tan, Xinyu Ma 等ACL 2025 · 被引用 19 次
- Reasons and Solutions for the Decline in Model Performance after EditingXiusheng Huang, Jiaxiang Liu, Yequan Wang, Kang LiuNeurIPS 2024 · 被引用 13 次
它引用的顶会 Paper15
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning 等ICML 2022 · 被引用 520 次
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 被引用 282 次
- Editable Neural NetworksAnton Sinitsin, Vsevolod Plokhotnyuk, Dmitry V. Pyrkin, Sergei Popov 等ICLR 2020 · 被引用 210 次
- Learn from Relational Correlations and Periodic Events for Temporal Knowledge Graph ReasoningKe Liang, Lingyuan Meng, Meng Liu, Yue Liu 等SIGIR 2023 · 被引用 117 次
相关 Paper
- EAMET: Robust Massive Model Editing via Embedding Alignment OptimizationYanbo Dai, Zhenlan Ji, Zongjie Li, Shuai WangICLR 2026
- Model Editing Harms General Abilities of Large Language Models: Regularization to the RescueJia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu 等EMNLP 2024 · 被引用 8 次
- Locate-then-edit for Multi-hop Factual Recall under Knowledge EditingZhuoran Zhang, Yongxiang Li, Zijian Kan, Keyuan Cheng 等ICML 2025
- SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding AlteringXiaopeng Li, Shasha Li, Shezheng Song, Huijun Liu 等AAAI 2025 · 被引用 11 次
- Parameter-Aware Contrastive Knowledge Editing: Tracing and Rectifying based on Critical Transmission PathsSonglin Zhai, Yuan Meng, Yuxin Zhang, Guilin QiACL 2025 · 被引用 3 次
