Evaluating Neuron Interpretation Methods of NLP Models
Yimin Fan, Fahim Dalvi, Nadir Durrani, Hassan Sajjad
Abstract
Neuron Interpretation has gained traction in the field of interpretability, and have provided fine-grained insights into what a model learns and how language knowledge is distributed amongst its different components. However, the lack of evaluation benchmark and metrics have led to siloed progress within these various methods, with very little work comparing them and highlighting their strengths and weaknesses. The reason for this discrepancy is the difficulty of creating ground truth datasets, for example, many neurons within a given model may learn the same phenomena, and hence there may not be one correct answer. Moreover, a learned phenomenon may spread across several neurons that work together -- surfacing these to create a gold standard challenging. In this work, we propose an evaluation framework that measures the compatibility of a neuron analysis method with other methods. We hypothesize that the more compatible a method is with the majority of the methods, the more confident one can be about its performance. We systematically evaluate our proposed framework and present a comparative analysis of a large set of neuron interpretation methods. We make the evaluation framework available to the community. It enables the evaluation of any new method using 20 concepts and across three pre-trained models.The code is released at https://github.com/fdalvi/neuron-comparative-analysis
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cd6c006-c673-43ec-90cd-eb4fc3ad73ebCited by top-tier papers4
- MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language ModelJiahao Huo, Yibo Yan, Boren Hu, Yutao Yue et al.EMNLP 2024 · 5 citations
- Discovering Influential Neuron Path in Vision TransformersYifan Wang, Yifei Liu, Yingdong Shi, Changming Li et al.ICLR 2025
- Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution MethodsHaeun Yu, Pepa Atanasova, Isabelle AugensteinACL 2024
- An Information-Theoretic Parameter-Free Bayesian Framework for Probing Labeled Dependency Trees from Attention ScoreHongxu Liu, Jing Ma, Xiaojie Wang, Caixia Yuan et al.ICLR 2026
Builds on15
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- When BERT Plays the Lottery, All Tickets Are WinningSai Prasanna, Anna Rogers, Anna RumshiskyEMNLP 2020 · 114 citations
- Discovering Latent Concepts Learned in BERTFahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani et al.ICLR 2022 · 74 citations
- On the Pitfalls of Analyzing Individual Neurons in Language ModelsOmer Antverg, Yonatan BelinkovICLR 2022 · 64 citations
Related papers
- Evaluating Neuron Explanations: A Unified Framework with Sanity ChecksTuomas P. Oikarinen, Ge Yan, Tsui-Wei WengICML 2025
- Interpretable Multi-dataset Evaluation for Named Entity RecognitionJinlan Fu, Pengfei Liu, Graham NeubigEMNLP 2020 · 49 citations
- RAVEL: Evaluating Interpretability Methods on Disentangling Language Model RepresentationsJing Huang, Zhengxuan Wu, Christopher Potts, Mor Geva et al.ACL 2024
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin et al.ACL 2025 · 8 citations
- Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description FrameworkLaura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer et al.NeurIPS 2025 · 12 citations
