(How) Can Transformers Predict Pseudo-Random Numbers?
Tao Tao, Darshil Doshi, Dayal Singh Kalra, Tianyu He, Maissam Barkeshli
摘要
Transformers excel at discovering patterns in sequential data, yet their fundamental limitations and learning mechanisms remain crucial topics of investigation. In this paper, we study the ability of Transformers to learn pseudo-random number sequences from linear congruential generators (LCGs), defined by the recurrence relation x t+1 = ax t + c mod m. We find that with sufficient architectural capacity and training data variety, Transformers can perform in-context prediction of LCG sequences with unseen moduli (m) and parameters (a, c). By analyzing the embedding layers and attention patterns, we uncover how Transformers develop algorithmic structures to learn these sequences in two scenarios of increasing complexity. First, we investigate how Transformers learn LCG sequences with unseen (a, c) but fixed modulus; and demonstrate successful learning up to m = 2 32 . We find that models learn to factorize m and utilize digit-wise number representations to make sequential predictions. In the second, more challenging scenario of unseen moduli, we show that Transformers can generalize to unseen moduli up to m test = 2 16 . In this case, the model employs a two-step strategy: first estimating the unknown modulus from the context, then utilizing prime factorizations to generate predictions. For this task, we observe a sharp transition in the accuracy at a critical depth d = 3. We also find that the number of in-context sequence elements needed to reach high accuracy scales sublinearly with the modulus. †
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- On The Geometry and Topology of Representations: the Manifolds of Modular AdditionGabriela Moisescu-Pareja, Gavin McCracken, Harley Wiltzer, Colin Daniels 等ICLR 2026 · 被引用 3 次
- Deep neural networks divide and conquer dihedral multiplicationSihui Wei, Gavin McCracken, Gabriela Moisescu-Pareja, Harley Wiltzer 等ICML 2026
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 被引用 324 次
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space SteeringSheng Liu, Haotian Ye, Lei Xing, James Y. ZouICML 2024 · 被引用 244 次
- Transformers Can Do Arithmetic with the Right EmbeddingsSean McLeish, Arpit Bansal, Alex Stein, Neel Jain 等NeurIPS 2024 · 被引用 94 次
- What learning algorithm is in-context learning? Investigations with linear modelsEkin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma 等ICLR 2023 · 被引用 85 次
相关 Paper
- Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and InterpretabilityTao Tao, Maissam BarkeshliICLR 2026
- Transformers in Pseudo-Random Number Generation: A Dual Perspective on Theory and PracticeRan Li, Lingshu ZengAAAI 2026
- Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable TasksRahul Ramesh, Ekdeep Singh Lubana, Mikail Khona, Robert P. Dick 等ICML 2024 · 被引用 17 次
- The Expressivity Limits of TransformersMaxime Meyer, Mario Michelessa, Caroline Chaux, Vincent TanICML 2026
- Looped Transformers for Length GeneralizationYing Fan, Yilun Du, Kannan Ramchandran, Kangwook LeeICLR 2025
