Disentangling Knowledge Representations for Large Language Model Editing
Mengqi Zhang, Zisheng Zhou, Xiaotian Ye, Qiang Liu, Zhaochun Ren, Zhumin Chen, Pengjie Ren
Abstract
Knowledge Editing has emerged as a promising solution for efficiently updating embedded knowledge in large language models (LLMs). While existing approaches demonstrate effectiveness in integrating new knowledge and preserving the original capabilities of LLMs, they fail to maintain fine-grained irrelevant knowledge, namely facts that share the same subject as edited knowledge but differ in relation and object. This challenge arises because subject representations inherently encode multiple attributes, causing the target and fine-grained irrelevant knowledge to become entangled in the representation space, and thus vulnerable to unintended alterations during editing. To address this, we propose DiKE, a novel approach that Disentangles Knowledge representations for LLM Editing (DiKE). DiKE consists of two key components: a Knowledge Representation Disentanglement (KRD) module that decomposes the subject representation into target-knowledge-related and -unrelated components, and a Disentanglementbased Knowledge Edit (DKE) module that updates only the target-related component while explicitly preserving the unrelated one. We further derive a closedform, rank-one parameter update based on matrix theory to enable efficient and minimally invasive edits. To rigorously evaluate fine-grained irrelevant knowledge preservation, we construct FINE-KED, a new benchmark comprising fine-grained irrelevant knowledge at different levels of relational similarity to the edited knowledge. Extensive experiments across multiple LLMs demonstrate that DiKE substantially improves fine-grained irrelevant knowledge preservation while maintaining competitive general editing performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83fddfc7-ed40-47b5-9353-81f70523483dCited by top-tier papers3
- LLM Unlearning Should Be Form-IndependentXiaotian Ye, Mengqi Zhang, Shu WuS&P 2026 · 3 citations
- Spectral Characterization and Mitigation of Sequential Knowledge Editing CollapseChi Zhang, Mengqi Zhang, Xiaotian Ye, Runxi Cheng et al.ACL 2026 · 2 citations
- StyliTruth : Unlocking Stylized yet Truthful LLM Generation via Disentangled SteeringChenglei Shen, Zhongxiang Sun, Teng Shi, Xiao Zhang et al.ICLR 2026
Builds on14
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
- Linearity of Relation Decoding in Transformer Language ModelsEvan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng et al.ICLR 2024 · 163 citations
- MELO: Enhancing Model Editing with Neuron-Indexed Dynamic LoRALang Yu, Qin Chen, Jie Zhou, Liang HeAAAI 2024 · 96 citations
- Massive Editing for Large Language Models via Meta LearningChenmien Tan, Ge Zhang, Jie FuICLR 2024 · 68 citations
Related papers
- AdaEdit: Advancing Continuous Knowledge Editing For Large Language ModelsQi Li, Xiaowen ChuACL 2025
- Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMsTingchao Fu, Wenkai Wang, Fanxiao Li, Huadong Zhang et al.ACL 2026
- Unveiling the Pitfalls of Knowledge Editing for Large Language ModelsZhoubo Li, Ningyu Zhang, Yunzhi Yao, Mengru Wang et al.ICLR 2024 · 47 citations
- GeoEdit: Geometric Knowledge Editing for Large Language ModelsYujie Feng, Li-Ming Zhan, Zexin Lu, Yongxin Xu et al.EMNLP 2025
- Conflict-Aware Knowledge Editing in the Wild: Semantic-Augmented Graph Representation for Unstructured TextZhange Zhang, Zhicheng Geng, Yuqing Ma, Tianbo Wang et al.NeurIPS 2025 · 2 citations
