Interpretability-based Tailored Knowledge Editing in Transformers
Yihuai Hong, Aldo Lipani
Abstract
Language models recognized as a new form of knowledge bases, face challenges of outdated, erroneous, and privacy-sensitive information, necessitating knowledge editing to rectify errors without costly retraining. Existing methods, spanning model's parameters modification, external knowledge integration, and in-context learning, lack in-depth analysis from a model interpretability perspective. Our work explores the instability in in-context learning outcomes, providing insights into its reasons and distinctions from other methods. Leveraging findings on the critical role of feed-forward MLPs in decoder-only models, we propose a tailored knowledge editing method, TailoredKE, that considers the unique information flow of each sample. Model interpretability reveals diverse attribute recall across transformer layers, guiding edits to specific features at different depths and mitigating over-editing issues.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58610b8b-cf78-42e5-a691-efe6a17552f6Cited by top-tier papers2
- TamEdit: Trajectory-Aware Meta-Learning for Specificity-Preserving Continual Knowledge EditingShiqiang Tian, Cheng Ding, Qin Chen, Jie Zhou et al.ACL 2026 · 1 citation
- GeoEdit: Geometric Knowledge Editing for Large Language ModelsYujie Feng, Li-Ming Zhan, Zexin Lu, Yongxin Xu et al.EMNLP 2025
Builds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
Related papers
- In-Context Editing: Learning Knowledge from Self-Induced DistributionsSiyuan Qi, Bangcheng Yang, Kailin Jiang, Xiaobo Wang et al.ICLR 2025
- Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMsJinzhe Liu, Junshu Sun, Shufan Shen, Chenxue Yang et al.NeurIPS 2025 · 8 citations
- Locate-then-edit for Multi-hop Factual Recall under Knowledge EditingZhuoran Zhang, Yongxiang Li, Zijian Kan, Keyuan Cheng et al.ICML 2025
- Everything is Editable: Extend Knowledge Editing to Unstructured Data in Large Language ModelsJingcheng Deng, Zihao Wei, Liang Pang, Hanxing Ding et al.ICLR 2025
- Can Knowledge be Transferred from Unimodal to Multimodal? Investigating the Transitivity of Multimodal Knowledge EditingLingyong Fang, Xinzhong Wang, Depeng Wang, Zongru Wu et al.ICCV 2025 · 4 citations
