Understanding and Relaxing the Limitations of Transformers for Linear Algebra
Andres Potapczynski, Alex Ali, Andrew Gordon Wilson
Abstract
Matrix operations, such as linear solves, eigendecompositions, and log determinants, are foundational building blocks for any number of downstream applications. Therefore, any broadly capable learning system should be able to effectively approximate these operations in its internal representation. Accordingly, there is great motivation to study transformers for linear algebra --- for if transformers cannot even semi-competently perform matrix operations, then we cannot expect them to form a basis for a generally intelligent system. We demonstrate that current techniques developing transformers for linear algebra have striking failure modes, prohibitive scaling, and particularly poor out-of-distribution generalization to other matrix distributions, and matrices of different sizes. Investigating further, we find that current transformer approaches operate as statistical interpolators, rather than discovering algorithms that will generalize to matrices from other distributions. Based on our understanding of these limitations, we develop a sequence of interventions that substantially improve scaling and performance, including matrix embeddings through a learnable projection, linear attention, looping, and a data pre-training distribution of structured matrices. We term the resulting method the RangeFormer, which we show has significantly improved scaling and performance on challenging OOD matrices from the matrix market. Moreover, with RangeFormer we show for the first time that transformers can be successfully applied to downstream tasks that involve iterative matrix operations, including Gaussian process learning, and improving the sampling distribution of randomized methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer et al.NeurIPS 2025 · 431 citations
- Looped Transformers as Programmable ComputersAngeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee et al.ICML 2023 · 175 citations
- Hungry Hungry Hippos: Towards Language Modeling with State Space ModelsDaniel Y. Fu, Tri Dao, Khaled Kamal Saab, Armin W. Thomas et al.ICLR 2023 · 117 citations
- Looped Transformers are Better at Learning Learning AlgorithmsLiu Yang, Kangwook Lee, Robert D. Nowak, Dimitris PapailiopoulosICLR 2024 · 82 citations
Related papers
- Towards Learning High-Precision Least Squares Algorithms with Sequence ModelsJerry Weihong Liu, Jessica Grogan, Owen M. Dugan, Ashish Rao et al.ICLR 2025
- FourierFormer: Transformer Meets Generalized Fourier Integral TheoremTan Nguyen, Minh Pham, Tam Nguyen, Khai Nguyen et al.NeurIPS 2022 · 59 citations
- Linear Transformers Implicitly Discover Unified Numerical AlgorithmsPatrick Lutz, Aditya Gangrade, Hadi Daneshmand, Venkatesh SaligramaNeurIPS 2025 · 3 citations
- Range-Invariant Approximation of Non-Linear Operations for Efficient BERT Fine-TuningJanghyeon Kim, Janghwan Lee, Jungwook Choi, JeongHo Han et al.DAC 2023 · 9 citations
- Neural Execution Engines: Learning to Execute SubroutinesYujun Yan, Kevin Swersky, Danai Koutra, Parthasarathy Ranganathan et al.NeurIPS 2020 · 47 citations
