Explainable and Efficient Editing for Large Language Models
Tianyu Zhang, Junfeng Fang, Houcheng Jiang, Baolong Bi, Xiang Wang, Xiangnan He
Abstract
Large Language Models (LLMs) exhibit remarkable capabilities in storing and retrieving vast amounts of factual knowledge. However, they retain outdated or incorrect information from Web corpora. Since full retraining is costly, locate-and-edit model editing methods offer a feasible alternative. Current methods typically follow a two-stage paradigm: (1) identifying critical layers that store knowledge and (2) updating their parameters to store new knowledge. However, both phases have their inherent limitations. Firstly, layer identification is independent of the knowledge being updated, ignoring the differences in knowledge storage patterns. Secondly, parameter updating suffers from high computational overhead due to gradient descent. To solve these, we propose an Explainable and effiCient model Editing method, termed ECE. Specifically, we integrate LLM explainability into the editing process, enabling the adaptive identification of the crucial neurons. Through clustering similar knowledge, we enable batch optimization in a single gradient step, significantly reducing computational time without compromising effectiveness. Extensive experiments demonstrate that ECE can achieve superior performance, showcasing the potential of explainability-driven editing methods for LLMs. Code is available at https://github.com/tianyuzhangterry/ECE . CCS Concepts • Computing methodologies → Semantic networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05b27fea-1e90-4f7e-9175-9a3ce2a96b6eCited by top-tier papers9
- Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language ModelsBaolong Bi, Shenghua Liu, Yiwei Wang, Yilong Xu et al.ICLR 2026 · 47 citations
- From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter EditingWei Liu, Hongkai Liu, Zhiying Deng, Yee-Whye Teh et al.ICML 2026 · 3 citations
- Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical AblationFrancesco Sovrano, Gabriele Dominici, Marc LangheinrichKDD 2026 · 3 citations
- DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal RegulationHoucheng Jiang, Zetong Zhao, Junfeng Fang, Haokai Ma et al.ICLR 2026 · 2 citations
- Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMsXinwei Wu, Heng Liu, Xiaohu Zhao, Yuqi Ren et al.AAAI 2026 · 2 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsTom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim et al.NeurIPS 2023 · 349 citations
Related papers
- Precise Localization of Memories: A Fine-grained Neuron-level Knowledge Editing Technique for LLMsHaowen Pan, Xiaozhi Wang, Yixin Cao, Zenglin Shi et al.ICLR 2025
- Fisher-Driven Adaptive Locating for Knowledge Editing in Large Language ModelsChenghao Xu, Jiexi Yan, Guangtao Lyu, Qi Liu et al.ACL 2026
- Locate-then-edit for Multi-hop Factual Recall under Knowledge EditingZhuoran Zhang, Yongxiang Li, Zijian Kan, Keyuan Cheng et al.ICML 2025
- Parameter-Aware Contrastive Knowledge Editing: Tracing and Rectifying based on Critical Transmission PathsSonglin Zhai, Yuan Meng, Yuxin Zhang, Guilin QiACL 2025 · 3 citations
- Scaling Knowledge Editing in LLMs to 100, 000 Facts with Neural KV DatabaseWeizhi Fei, Hao Shi, Jing Xu, Jingchen Peng et al.ICLR 2026 · 2 citations
