SAME: Safety-Aware Model Editing Guided by Safety Transformation
Jiayi Wang, Shipeng Wang, Ji Wu, Jian Sun
摘要
Editing large language models is challenging as incorporating new knowledge often requires sequential parameter updates while maintaining model capability. In this work, we experimentally observe that sequential knowledge updating under the locate-then-edit framework can introduce safety risks, regardless of whether the knowledge being edited is benign or malicious. We propose a novel model editing approach that estimates safety transforms and identifies corresponding safety direction in the neural activation space, and then aligns neural activation updates and network parameter updates under the safety constraints, resulting in a safety-aware model editing approach. We evaluate our approach on open-source LLMs, Llama-3-8B-Instruct, Qwen3-4B-Instruct and Qwen2.5-14B-Instruct, using the benchmark datasets ZsRE and COUNTERFACT, as well as the malicious dataset Mal-KSet. Experimental results demonstrate that our approach effectively reduces unsafe responses to malicious queries while preserving the effectiveness of model editing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn 等ICLR 2022 · 被引用 527 次
- Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsTom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim 等NeurIPS 2023 · 被引用 349 次
相关 Paper
- Scaling Knowledge Editing in LLMs to 100, 000 Facts with Neural KV DatabaseWeizhi Fei, Hao Shi, Jing Xu, Jingchen Peng 等ICLR 2026 · 被引用 2 次
- AlphaEdit: Null-Space Constrained Knowledge Editing for Language ModelsJunfeng Fang, Houcheng Jiang, Kun Wang, Yunshan Ma 等ICLR 2025 · 被引用 1 次
- Should We Really Edit Language Models? On the Evaluation of Edited Language ModelsQi Li, Xiang Liu, Zhenheng Tang, Peijie Dong 等NeurIPS 2024 · 被引用 25 次
- Multilingual Safety Alignment Via Sparse Weight EditingJiaming Liang, Zhaoxin Wang, Handing WangICML 2026 · 被引用 3 次
- Diagnosing Hidden Instabilities in Model Editing via Uncertainty QuantificationZihan Gu, Tianyi Zhang, Xinyan Zhang, Zhiyuan Wang 等ACL 2026
