On the Pitfalls of Analyzing Individual Neurons in Language Models
Omer Antverg, Yonatan Belinkov
Abstract
While many studies have shown that linguistic information is encoded in hidden word representations, few have studied individual neurons, to show how and in which neurons it is encoded. Among these, the common approach is to use an external probe to rank neurons according to their relevance to some linguistic attribute, and to evaluate the obtained ranking using the same probe that produced it. We show two pitfalls in this methodology: 1. It confounds distinct factors: probe quality and ranking quality. We separate them and draw conclusions on each. 2. It focuses on encoded information, rather than information that is used by the model. We show that these are not the same. We compare two recent ranking methods and a simple one we introduce, and evaluate them with regard to both of these aspects. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc4f4b0f-5cdd-4358-8f39-839f40484496Cited by top-tier papers21
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsBoyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie et al.ICML 2024 · 215 citations
- Confidence Regulation Neurons in Language ModelsAlessandro Stolfo, Ben Wu, Wes Gurnee, Yonatan Belinkov et al.NeurIPS 2024 · 68 citations
- Decomposing and Editing Predictions by Modeling Model ComputationHarshay Shah, Andrew Ilyas, Aleksander MadryICML 2024 · 25 citations
- NeuroStrike: Neuron-Level Attacks on Aligned LLMsLichao Wu, Sasha Behrouzi, Mohamadreza Rostami, Maximilian Thang et al.NDSS 2026 · 21 citations
- Finding Skill Neurons in Pre-trained Transformer-based Language ModelsXiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou et al.EMNLP 2022 · 19 citations
Builds on3
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Analyzing Individual Neurons in Pre-trained Language ModelsNadir Durrani, Hassan Sajjad, Fahim Dalvi, Yonatan BelinkovEMNLP 2020 · 5 citations
- Intrinsic Probing through Dimension SelectionLucas Torroba Hennigen, Adina Williams, Ryan CotterellEMNLP 2020 · 3 citations
Related papers
- A Latent-Variable Model for Intrinsic ProbingKarolina Stanczak, Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell et al.AAAI 2023 · 6 citations
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 34 citations
- Probing Linguistic Features of Sentence-Level Representations in Relation ExtractionChristoph Alt, Aleksandra Gabryszak, Leonhard HennigACL 2020 · 29 citations
- Information-Theoretic Probing for Linguistic StructureTiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod et al.ACL 2020 · 21 citations
- Exploring the Role of BERT Token Representations to Explain Sentence Probing ResultsHosein Mohebbi, Ali Modarressi, Mohammad Taher PilehvarEMNLP 2021 · 14 citations
