Neuron Empirical Gradient: Discovering and Quantifying Neurons' Global Linear Controllability
Xin Zhao, Zehui Jiang, Naoki Yoshinaga
Abstract
While feed-forward neurons in pre-trained language models (PLMs) can encode knowledge, past research targeted a small subset of neurons that heavily influence outputs. This leaves the broader role of neuron activations unclear, limiting progress in areas like knowledge editing. We uncover a global linear relationship between neuron activations and outputs using neuron interventions on a knowledge probing dataset. The gradient of this linear relationship, which we call the neuron empirical gradient (NEG), captures how changes in activations affect predictions. To compute NEG efficiently, we propose NeurGrad, enabling large-scale analysis of neuron behavior in PLMs. We also show that NEG effectively captures language skills across diverse prompts through skill neuron probing. Experiments on MCEval8k, a multi-genre multiple-choice knowledge benchmark, support NEG's ability to represent model knowledge. Further analysis highlights the key properties of NEG-based skill representation: efficiency, robustness, flexibility, and interdependency. The code and data are released. xzhao-tkl/NEG iszhaoxin/MCEval8K !"#$%&'()'(%)*&+,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b744caed-8ecd-4046-9020-67f0b93ff7baCited by top-tier papers1
Ask how each one uses itBuilds on17
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie et al.EMNLP 2023 · 224 citations
- A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityAndrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg et al.ICML 2024 · 177 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
Related papers
- Towards Neuron Attributions in Multi-Modal Large Language ModelsJunfeng Fang, Zac Bi, Ruipeng Wang, Houcheng Jiang et al.NeurIPS 2024 · 16 citations
- Journey to the Center of the Knowledge Neurons: Discoveries of Language-Independent Knowledge Neurons and Degenerate Knowledge NeuronsYuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu et al.AAAI 2024 · 64 citations
- Neuron-Level Analysis of Cultural Understanding in Large Language ModelsTaisei Yamamoto, Ryoma Kumon, Danushka Bollegala, Hitomi YanakaICLR 2026 · 1 citation
- Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMsJinzhe Liu, Junshu Sun, Shufan Shen, Chenxue Yang et al.NeurIPS 2025 · 8 citations
- Does Large Language Model Contain Task-Specific Neurons?Ran Song, Shizhu He, Shuting Jiang, Yantuan Xian et al.EMNLP 2024 · 1 citation
