Trapped by simplicity: When Transformers fail to learn from noisy features
Evan Peters, Matheus Hrabowec Zambianco, Ando Deng, Devin Blankespoor, Achim Kempf
Abstract
Noise is ubiquitous in data used to train large language models, but it is not well understood whether these models are able to correctly generalize to inputs generated without noise. Here, we study noise-robust learning: are transformers trained on data with noisy features able to find a target function that correctly predicts labels for noiseless features? We show that transformers succeed at noise-robust learning for a selection of k-sparse parity and majority functions, compared to LSTMs which fail at this task for even modest feature noise. However, we find that transformers typically fail at noise-robust learning of random k-juntas, especially when the boolean sensitivity of the optimal solution is smaller than that of the target function. We argue that this failure is due to a combination of two factors: transformers' bias toward simpler functions, combined with an observation that the optimal function for noise-robust learning typically has lower sensitivity than the target function for random boolean functions. We test this hypothesis by exploiting transformers' simplicity bias to trap them in an incorrect solution, but show that transformers can escape this trap by training with an additional loss term penalizing high-sensitivity solutions. Overall, we find that transformers are particularly ineffective for learning boolean functions in the presence of feature noise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 866f5cd4-2922-4111-8c31-2b2983bfc883Builds on13
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- Towards Understanding Grokking: An Effective Theory of Representation LearningZiming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud et al.NeurIPS 2022 · 299 citations
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt et al.ICLR 2024 · 243 citations
- Understanding The Robustness in Vision TransformersDaquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao et al.ICML 2022 · 242 citations
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade et al.NeurIPS 2022 · 220 citations
Related papers
- Simplicity Bias in Transformers and their Ability to Learn Sparse Boolean FunctionsSatwik Bhattamishra, Arkil Patel, Varun Kanade, Phil BlunsomACL 2023 · 7 citations
- Why are Sensitive Functions Hard for Transformers?Michael Hahn, Mark RofinACL 2024 · 3 citations
- Understanding the Parameter Space Geometry of Transformers Encoding Boolean FunctionsBlanka Kövér, Alexandra Butoi, Anej Svete, Michael Hahn et al.ICML 2026
- Why Larger Language Models Do In-context Learning Differently?Zhenmei Shi, Junyi Wei, Zhuoyan Xu, Yingyu LiangICML 2024 · 54 citations
- Transformers Learn Low Sensitivity Functions: Investigations and ImplicationsBhavya Vasudeva, Deqing Fu, Tianyi Zhou, Elliott Kau et al.ICLR 2025
