Evaluating Neuron Interpretation Methods of NLP Models
Yimin Fan, Fahim Dalvi, Nadir Durrani, Hassan Sajjad
摘要
Neuron Interpretation has gained traction in the field of interpretability, and have provided fine-grained insights into what a model learns and how language knowledge is distributed amongst its different components. However, the lack of evaluation benchmark and metrics have led to siloed progress within these various methods, with very little work comparing them and highlighting their strengths and weaknesses. The reason for this discrepancy is the difficulty of creating ground truth datasets, for example, many neurons within a given model may learn the same phenomena, and hence there may not be one correct answer. Moreover, a learned phenomenon may spread across several neurons that work together -- surfacing these to create a gold standard challenging. In this work, we propose an evaluation framework that measures the compatibility of a neuron analysis method with other methods. We hypothesize that the more compatible a method is with the majority of the methods, the more confident one can be about its performance. We systematically evaluate our proposed framework and present a comparative analysis of a large set of neuron interpretation methods. We make the evaluation framework available to the community. It enables the evaluation of any new method using 20 concepts and across three pre-trained models.The code is released at https://github.com/fdalvi/neuron-comparative-analysis
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language ModelJiahao Huo, Yibo Yan, Boren Hu, Yutao Yue 等EMNLP 2024 · 被引用 5 次
- Discovering Influential Neuron Path in Vision TransformersYifan Wang, Yifei Liu, Yingdong Shi, Changming Li 等ICLR 2025
- Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution MethodsHaeun Yu, Pepa Atanasova, Isabelle AugensteinACL 2024
- An Information-Theoretic Parameter-Free Bayesian Framework for Probing Labeled Dependency Trees from Attention ScoreHongxu Liu, Jing Ma, Xiaojie Wang, Caixia Yuan 等ICLR 2026
它引用的顶会 Paper15
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 被引用 229 次
- When BERT Plays the Lottery, All Tickets Are WinningSai Prasanna, Anna Rogers, Anna RumshiskyEMNLP 2020 · 被引用 114 次
- Discovering Latent Concepts Learned in BERTFahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani 等ICLR 2022 · 被引用 74 次
- On the Pitfalls of Analyzing Individual Neurons in Language ModelsOmer Antverg, Yonatan BelinkovICLR 2022 · 被引用 64 次
相关 Paper
- Evaluating Neuron Explanations: A Unified Framework with Sanity ChecksTuomas P. Oikarinen, Ge Yan, Tsui-Wei WengICML 2025
- Interpretable Multi-dataset Evaluation for Named Entity RecognitionJinlan Fu, Pengfei Liu, Graham NeubigEMNLP 2020 · 被引用 49 次
- RAVEL: Evaluating Interpretability Methods on Disentangling Language Model RepresentationsJing Huang, Zhengxuan Wu, Christopher Potts, Mor Geva 等ACL 2024
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin 等ACL 2025 · 被引用 8 次
- Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description FrameworkLaura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer 等NeurIPS 2025 · 被引用 12 次
