Evaluating Neuron Explanations: A Unified Framework with Sanity Checks
Tuomas P. Oikarinen, Ge Yan, Tsui-Wei Weng
Abstract
Understanding the function of individual units in a neural network is an important building block for mechanistic interpretability. This is often done by generating a simple text explanation of the behavior of individual neurons or units. For these explanations to be useful, we must understand how reliable and truthful they are. In this work we unify many existing explanation evaluation methods under one mathematical framework. This allows us to compare existing evaluation metrics, understand the evaluation pipeline with increased clarity and apply existing statistical methods on the evaluation. In addition, we propose two simple sanity checks on the evaluation metrics and show that many commonly used metrics fail these tests and do not change their score after massive changes to the concept labels. Based on our experimental and theoretical results, we propose guidelines that future evaluations should follow and identify a set of reliable evaluation metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b6e108a-4add-4fd9-9037-4cc3557eed67Cited by top-tier papers4
- CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder FeaturesSeonglae Cho, Zekun Wu, Adriano KoshiyamaICML 2026 · 5 citations
- Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated InterpretabilityTuomas Oikarinen, Ge Yan, Akshay Kulkarni, Tsui-Wei WengCVPR 2026 · 1 citation
- Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse AutoencodersXiangchen Song, Aashiq Muhamed, Yujia Zheng, Lingjing Kong et al.ACL 2026
- ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning ModelsChung-En Sun, Ge Yan, Tsui-Wei WengEMNLP 2025
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
Related papers
- CoSy: Evaluating Textual Explanations of NeuronsLaura Kopf, Philine Lou Bommer, Anna Hedström, Sebastian Lapuschkin et al.NeurIPS 2024 · 23 citations
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin et al.ACL 2025 · 8 citations
- Evaluating the Robustness of Interpretability Methods through Explanation Invariance and EquivarianceJonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 27 citations
- Evaluating Explanation Methods for Neural Machine TranslationJierui Li, Lemao Liu, Huayang Li, Guanlin Li et al.ACL 2020 · 24 citations
- Evaluating Neuron Interpretation Methods of NLP ModelsYimin Fan, Fahim Dalvi, Nadir Durrani, Hassan SajjadNeurIPS 2023 · 11 citations
