Confidence Regulation Neurons in Language Models
Alessandro Stolfo, Ben Wu, Wes Gurnee, Yonatan Belinkov, Xingyi Song, Mrinmaya Sachan, Neel Nanda
Abstract
Despite their widespread use, the mechanisms by which large language models (LLMs) represent and regulate uncertainty in next-token predictions remain largely unexplored. This study investigates two critical components believed to influence this uncertainty: the recently discovered entropy neurons and a new set of components that we term token frequency neurons. Entropy neurons are characterized by an unusually high weight norm and influence the final layer normalization (LayerNorm) scale to effectively scale down the logits. Our work shows that entropy neurons operate by writing onto an unembedding null space, allowing them to impact the residual stream norm with minimal direct effect on the logits themselves. We observe the presence of entropy neurons across a range of models, up to 7 billion parameters. On the other hand, token frequency neurons, which we discover and describe here for the first time, boost or suppress each token's logit proportionally to its log frequency, thereby shifting the output distribution towards or away from the unigram distribution. Finally, we present a detailed case study where entropy neurons actively manage confidence in the setting of induction, i.e. detecting and continuing repeated subsequences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext adfaad16-eab1-4915-aa1b-bfe6143fee1aCited by top-tier papers19
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety NeuronsJianhui Chen, Xiaozhi Wang, Zijun Yao, Yushi Bai et al.NeurIPS 2025 · 53 citations
- Causal Sufficiency and Necessity Improves Chain-of-Thought ReasoningXiangning Yu, Zhuohan Wang, Linyi Yang, Haoxuan Li et al.NeurIPS 2025 · 19 citations
- Dense SAE Latents Are Features, Not BugsXiaoqing Sun, Alessandro Stolfo, Joshua Engels, Ben Wu et al.NeurIPS 2025 · 19 citations
- Transferring Linear Features Across Language Models With Model StitchingAlan Chen, Jack Merullo, Alessandro Stolfo, Ellie PavlickNeurIPS 2025 · 17 citations
- How do LLMs Compute Verbal Confidence?Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero et al.ICML 2026 · 16 citations
Builds on22
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
Related papers
- Forget What Matters, Keep the Rest: Selective Unlearning of Informative TokensSeunghee Koh, Sunghyun Baek, Youngdong Kim, Junmo KimACL 2026 · 1 citation
- Interpreting Context Look-ups in Transformers: Investigating Attention-MLP InteractionsClement Neo, Shay B. Cohen, Fazl BarezEMNLP 2024 · 3 citations
- Calibration Across Layers: Understanding Calibration Evolution in LLMsAbhinav Joshi, Areeb Ahmad, Ashutosh ModiEMNLP 2025 · 1 citation
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-trainingWoojin Chung, Jeonghoon KimNeurIPS 2025 · 6 citations
- Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic InterpretabilityLuca Baroni, Galvin Khara, Joachim Schaeffer, Marat Subkhankulov et al.ICLR 2026 · 8 citations
