What does the Knowledge Neuron Thesis Have to do with Knowledge?
Jingcheng Niu, Andrew Liu, Zining Zhu, Gerald Penn
Abstract
We reassess the Knowledge Neuron (KN) Thesis: an interpretation of the mechanism underlying the ability of large language models to recall facts from a training corpus. This nascent thesis proposes that facts are recalled from the training corpus through the MLP weights in a manner resembling key-value memory, implying in effect that "knowledge" is stored in the network. Furthermore, by modifying the MLP modules, one can control the language model's generation of factual information. The plausibility of the KN thesis has been demonstrated by the success of KN-inspired model editing methods (Dai et al., 2022; Meng et al., 2022) . We find that this thesis is, at best, an oversimplification. Not only have we found that we can edit the expression of certain linguistic phenomena using the same model editing methods but, through a more comprehensive evaluation, we have found that the KN thesis does not adequately explain the process of factual expression. While it is possible to argue that the MLP weights store complex patterns that are interpretable both syntactically and semantically, these patterns do not constitute "knowledge." To gain a more comprehensive understanding of the knowledge representation process, we must look beyond the MLP weights and explore recent models' complex layer structures and attention mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70ee5bdb-9169-4c2c-9086-0ff307bad462Cited by top-tier papers23
- WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language ModelsPeng Wang, Zexi Li, Ningyu Zhang, Ziwen Xu et al.NeurIPS 2024 · 125 citations
- Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic ModelsAviv Bick, Kevin Y. Li, Eric P. Xing, J. Zico Kolter et al.NeurIPS 2024 · 78 citations
- Knowledge Circuits in Pretrained TransformersYunzhi Yao, Ningyu Zhang, Zekun Xi, Mengru Wang et al.NeurIPS 2024 · 71 citations
- NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-TuningXin Yi, Shunfan Zheng, Linlin Wang, Gerard de Melo et al.AAAI 2025 · 38 citations
- Can Editing LLMs Inject Harm?Canyu Chen, Baixiang Huang, Zekun Li, Zhaorun Chen et al.AAAI 2026 · 26 citations
Builds on13
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
- Interpretability at Scale: Identifying Causal Mechanisms in AlpacaZhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts et al.NeurIPS 2023 · 146 citations
- Editing Large Language Models: Problems, Methods, and OpportunitiesYunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng et al.EMNLP 2023 · 83 citations
Related papers
- Knowledge Localization: Mission Not Accomplished? Enter Query Localization!Yuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu et al.ICLR 2025
- Cracking Factual Knowledge: A Comprehensive Analysis of Degenerate Knowledge Neurons in Large Language ModelsYuheng Chen, Pengfei Cao, Yubo Chen, Yining Wang et al.ACL 2025
- Knowledge Neurons in Pretrained TransformersDamai Dai, Li Dong, Yaru Hao, Zhifang Sui et al.ACL 2022
- Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsPeter Hase, Mohit Bansal, Been Kim, Asma GhandehariounNeurIPS 2023 · 307 citations
- Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic LocalizationPhillip Guo, Aaquib Syed, Abhay Sheshadri, Aidan Ewart et al.ICML 2025
