Learning from Emergence: A Study on Proactively Inhibiting the Monosemantic Neurons of Artificial Neural Networks
Jiachuan Wang, Shimin Di, Lei Chen, Charles Wang Wai Ng
摘要
Recently, emergence has received widespread attention from the research community along with the success of large-scale models. Different from the literature, we hypothesize a key factor that promotes the performance during the increase of scale: the reduction of monosemantic neurons that can only form one-to-one correlations with specific features. Monosemantic neurons tend to be sparser and have negative impacts on the performance in large models. Inspired by this insight, we propose an intuitive idea to identify monosemantic neurons and inhibit them. However, achieving this goal is a non-trivial task as there is no unified quantitative evaluation metric and simply banning monosemantic neurons does not promote polysemanticity in neural networks. Therefore, we first propose a new metric to measure the monosemanticity of neurons with the guarantee of efficiency for online computation, then introduce a theoretically supported method to suppress monosemantic neurons and proactively promote the ratios of polysemantic neurons in training neural networks. We validate our conjecture that monosemanticity brings about performance change at different model scales on a variety of neural networks and benchmark datasets in different areas, including language, image, and physics simulation tasks. Further experiments validate our analysis and theory regarding the inhibition of monosemanticity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted RemedyJingyi Cui, Qi Zhang, Yifei Wang, Yisen WangICLR 2026 · 被引用 16 次
- Measuring and Guiding MonosemanticityRuben Härle, Felix Friedrich, Manuel Brack, Björn Deiseroth 等NeurIPS 2025 · 被引用 12 次
- Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model InfluenceBofan Gong, Shiyang Lai, James Evans, Dawn SongICLR 2026 · 被引用 4 次
- Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation PerspectiveHanqi Yan, Yanzheng Xiang, Guangyi Chen, Yifei Wang 等EMNLP 2024 · 被引用 1 次
- TGDD: Trajectory Guided Dataset Distillation with Balanced DistributionFengli Ran, Xiao Pu, Bo Liu, Xiuli Bi 等AAAI 2026
它引用的顶会 Paper5
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 被引用 796 次
- Analyzing Transformers in Embedding SpaceGuy Dar, Mor Geva, Ankit Gupta, Jonathan BerantACL 2023 · 被引用 36 次
- Transformer Feed-Forward Layers Are Key-Value MemoriesMor Geva, Roei Schuster, Jonathan Berant, Omer LevyEMNLP 2021 · 被引用 33 次
相关 Paper
- Expand Neurons, Not ParametersLinghao Kong, Inimai Subramanian, Yonadav Shavit, Micah Adler 等ICML 2026 · 被引用 1 次
- The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert LevelJeremy Herbst, Stefan Wermter, Jae Hee LeeICML 2026 · 被引用 9 次
- Wasserstein Distances, Neuronal Entanglement, and SparsityShashata Sawmya, Linghao Kong, Ilia Markov, Dan Alistarh 等ICLR 2025
- Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous WordsGouki Minegishi, Hiroki Furuta, Yusuke Iwasawa, Yutaka MatsuoICLR 2025
- Quantifying Semantic Emergence in Language ModelsHang Chen, Xinyu Yang, Jiaying Zhu, Wenya WangACL 2025
