Explainable and Efficient Editing for Large Language Models
Tianyu Zhang, Junfeng Fang, Houcheng Jiang, Baolong Bi, Xiang Wang, Xiangnan He
摘要
Large Language Models (LLMs) exhibit remarkable capabilities in storing and retrieving vast amounts of factual knowledge. However, they retain outdated or incorrect information from Web corpora. Since full retraining is costly, locate-and-edit model editing methods offer a feasible alternative. Current methods typically follow a two-stage paradigm: (1) identifying critical layers that store knowledge and (2) updating their parameters to store new knowledge. However, both phases have their inherent limitations. Firstly, layer identification is independent of the knowledge being updated, ignoring the differences in knowledge storage patterns. Secondly, parameter updating suffers from high computational overhead due to gradient descent. To solve these, we propose an Explainable and effiCient model Editing method, termed ECE. Specifically, we integrate LLM explainability into the editing process, enabling the adaptive identification of the crucial neurons. Through clustering similar knowledge, we enable batch optimization in a single gradient step, significantly reducing computational time without compromising effectiveness. Extensive experiments demonstrate that ECE can achieve superior performance, showcasing the potential of explainability-driven editing methods for LLMs. Code is available at https://github.com/tianyuzhangterry/ECE . CCS Concepts • Computing methodologies → Semantic networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language ModelsBaolong Bi, Shenghua Liu, Yiwei Wang, Yilong Xu 等ICLR 2026 · 被引用 47 次
- From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter EditingWei Liu, Hongkai Liu, Zhiying Deng, Yee-Whye Teh 等ICML 2026 · 被引用 3 次
- Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical AblationFrancesco Sovrano, Gabriele Dominici, Marc LangheinrichKDD 2026 · 被引用 3 次
- DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal RegulationHoucheng Jiang, Zetong Zhao, Junfeng Fang, Haokai Ma 等ICLR 2026 · 被引用 2 次
- Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMsXinwei Wu, Heng Liu, Xiaohu Zhao, Yuqi Ren 等AAAI 2026 · 被引用 2 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsTom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim 等NeurIPS 2023 · 被引用 349 次
相关 Paper
- Precise Localization of Memories: A Fine-grained Neuron-level Knowledge Editing Technique for LLMsHaowen Pan, Xiaozhi Wang, Yixin Cao, Zenglin Shi 等ICLR 2025
- Fisher-Driven Adaptive Locating for Knowledge Editing in Large Language ModelsChenghao Xu, Jiexi Yan, Guangtao Lyu, Qi Liu 等ACL 2026
- Locate-then-edit for Multi-hop Factual Recall under Knowledge EditingZhuoran Zhang, Yongxiang Li, Zijian Kan, Keyuan Cheng 等ICML 2025
- Parameter-Aware Contrastive Knowledge Editing: Tracing and Rectifying based on Critical Transmission PathsSonglin Zhai, Yuan Meng, Yuxin Zhang, Guilin QiACL 2025 · 被引用 3 次
- Scaling Knowledge Editing in LLMs to 100, 000 Facts with Neural KV DatabaseWeizhi Fei, Hao Shi, Jing Xu, Jingchen Peng 等ICLR 2026 · 被引用 2 次
