Why are Sensitive Functions Hard for Transformers?
Michael Hahn, Mark Rofin
Abstract
Empirical studies have identified a range of learnability biases and limitations of transformers, such as a persistent difficulty in learning to compute simple formal languages such as PARITY, and a bias towards low-degree functions. However, theoretical understanding remains limited, with existing expressiveness theory either overpredicting or underpredicting realistic learning abilities. We prove that, under the transformer architecture, the loss landscape is constrained by the inputspace sensitivity: Transformers whose output is sensitive to many parts of the input string inhabit isolated points in parameter space, leading to a low-sensitivity bias in generalization. We show theoretically and empirically that this theory unifies a broad array of empirical observations about the learning abilities and biases of transformers, such as their generalization bias towards low sensitivity and low degree, and difficulty in length generalization for PARITY. This shows that understanding transformers' inductive biases requires studying not just their in-principle expressivity, but also their loss landscape.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers35
- The Expressive Capacity of State Space Models: A Formal Language PerspectiveYash Raj Sarrof, Yana Veitsman, Michael HahnNeurIPS 2024 · 53 citations
- How Far Can Transformers Reason? The Globality Barrier and Inductive ScratchpadEmmanuel Abbe, Samy Bengio, Aryo Lotfi, Colin Sandon et al.NeurIPS 2024 · 52 citations
- ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMsLandon Butler, Abhineet Agarwal, Justin Singh Kang, Yigit Efe Erginbas et al.NeurIPS 2025 · 19 citations
- The Expressive Limits of Diagonal SSMs for State-TrackingMehran Shakerinava, Behnoush Khavari, Siamak Ravanbakhsh, Sarath ChandarICLR 2026 · 11 citations
- Lost in Transmission: When and Why LLMs Fail to Reason GloballyTobias Schnabel, Kiran Tomlinson, Adith Swaminathan, Jennifer NevilleNeurIPS 2025 · 9 citations
Builds on21
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
- Towards Revealing the Mystery behind Chain of Thought: A Theoretical PerspectiveGuhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye et al.NeurIPS 2023 · 470 citations
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz et al.NeurIPS 2022 · 267 citations
Related papers
- Understanding the Parameter Space Geometry of Transformers Encoding Boolean FunctionsBlanka Kövér, Alexandra Butoi, Anej Svete, Michael Hahn et al.ICML 2026
- Simplicity Bias in Transformers and their Ability to Learn Sparse Boolean FunctionsSatwik Bhattamishra, Arkil Patel, Varun Kanade, Phil BlunsomACL 2023 · 7 citations
- Transformers Learn Low Sensitivity Functions: Investigations and ImplicationsBhavya Vasudeva, Deqing Fu, Tianyi Zhou, Elliott Kau et al.ICLR 2025
- Towards Understanding Inductive Bias in Transformers: A View From InfinityItay Lavie, Guy Gur-Ari, Zohar RingelICML 2024 · 11 citations
- Trapped by simplicity: When Transformers fail to learn from noisy featuresEvan Peters, Matheus Hrabowec Zambianco, Ando Deng, Devin Blankespoor et al.ICLR 2026
