SAKE: Steering Activations for Knowledge Editing
Marco Scialanga, Thibault Laugel, Vincent Grari, Marcin Detyniecki
Abstract
As Large Langue Models have been shown to memorize real-world facts, the need to update this knowledge in a controlled and efficient manner arises. Designed with these constraints in mind, Knowledge Editing (KE) approaches propose to alter specific facts in pretrained models. However, they have been shown to suffer from several limitations, including their lack of contextual robustness and their failure to generalize to logical implications related to the fact. To overcome these issues, we propose SAKE, a steering activation method that models a fact to be edited as a distribution rather than a single prompt. Leveraging Optimal Transport, SAKE alters the LLM behavior over a whole fact-related distribution, defined as paraphrases and logical implications. Several numerical experiments demonstrate the effectiveness of this method: SAKE is thus able to perform more robust edits than its existing counterparts. The code to reproduce all experiments is made available on a repository 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd889166-c347-4a97-9b6b-49a74b24f318Cited by top-tier papers1
Ask how each one uses itBuilds on16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning et al.ICML 2022 · 520 citations
- Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsTom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim et al.NeurIPS 2023 · 349 citations
Related papers
- MoEEdit: Efficient and Routing-Stable Knowledge Editing for Mixture-of-Experts LLMsYupu Gu, Rongzhe Wei, Andy Zhu, Pan LiICLR 2026 · 4 citations
- Serial Lifelong Editing via Mixture of Knowledge ExpertsYuJu Cheng, Yu-Chu Yu, Kai-Po Chang, Yu-Chiang Frank WangACL 2025
- Can We Edit Factual Knowledge by In-Context Learning?Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan et al.EMNLP 2023 · 40 citations
- REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge EditingHaitian Zhong, Yuhuan Liu, Ziyang Xu, Guofan Liu et al.EMNLP 2025
- FAME: Towards Factual Multi-Task Model EditingZeng Li, Yingyu Shan, Zeming Liu, Jiashu Yao et al.EMNLP 2024 · 1 citation
