Noise Stability of Transformer Models
Themistoklis Haris, Zihan Zhang, Yuichi Yoshida
Abstract
Understanding simplicity biases in deep learning offers a promising path toward developing reliable AI. A common metric for this, inspired by Boolean function analysis, is average sensitivity, which captures a model's robustness to singletoken perturbations. We argue that average sensitivity has two key limitations: it lacks a natural generalization to real-valued domains and fails to explain the "junta-like" input dependence we empirically observe in modern LLMs. To address these limitations, we propose noise stability as a more comprehensive simplicity metric. Noise stability expresses a model's robustness to correlated noise applied to all input coordinates simultaneously. We provide a theoretical analysis of noise stability for single-layer attention and ReLU MLP layers and tackle the multi-layer propagation problem with a covariance interval propagation approach. Building on this theory, we develop a practical noise stability regularization method. Experiments on algorithmic and next-token-prediction tasks show that our regularizer consistently catalyzes grokking and accelerates training by approximately 35% and 75% respectively. Our results sculpt a new connection between signal propagation in neural networks and interpretability, with noise stability emerging as a powerful tool for understanding and improving modern Transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3109d58a-1828-4a02-9963-ec3ea708e07cBuilds on18
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain et al.NeurIPS 2020 · 503 citations
- A Signal Propagation Perspective for Pruning Neural Networks at InitializationNamhoon Lee, Thalaiyasingam Ajanthan, Stephen Gould, Philip H. S. TorrICLR 2020 · 174 citations
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank CollapseLorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto et al.NeurIPS 2022 · 161 citations
- A unifying view on implicit bias in training linear neural networksChulhee Yun, Shankar Krishnan, Hossein MobahiICLR 2021 · 94 citations
Related papers
- Transformers Learn Low Sensitivity Functions: Investigations and ImplicationsBhavya Vasudeva, Deqing Fu, Tianyi Zhou, Elliott Kau et al.ICLR 2025
- Trapped by simplicity: When Transformers fail to learn from noisy featuresEvan Peters, Matheus Hrabowec Zambianco, Ando Deng, Devin Blankespoor et al.ICLR 2026
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon et al.ICML 2022 · 144 citations
- Random Spiking Neural Networks are Stable and Spectrally SimpleErnesto Araya, Massimiliano Datres, Gitta KutyniokICLR 2026 · 1 citation
- Noise Injection Node Regularization for Robust LearningNoam Levi, Itay M. Bloch, Marat Freytsis, Tomer VolanskyICLR 2023 · 2 citations
