Neural Networks Learn Statistics of Increasing Complexity
Nora Belrose, Quintin Pope, Lucia Quirke, Alex Mallen, Xiaoli Z. Fern
Abstract
The distributional simplicity bias (DSB) posits that neural networks learn low-order moments of the data distribution first, before moving on to higher-order correlations. In this work, we present compelling new evidence for the DSB by showing that networks automatically learn to perform well on maximum-entropy distributions whose low-order statistics match those of the training set early in training, then lose this ability later. We also extend the DSB to discrete domains by proving an equivalence between token n-gram frequencies and the moments of embedding vectors, and by finding empirical evidence for the bias in LLMs. Finally we use optimal transport methods to surgically edit the low-order statistics of one class of images to match those of another, and show early-training networks treat the edited images as if they were drawn from the target class. Code is available at https://github.com/ EleutherAI/features-across-time .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9eed23b-0352-43e0-b0ac-113d21ce6ebcCited by top-tier papers16
- Tracing the Representation Geometry of Language Models from Pretraining to Post-trainingMelody Zixuan Li, Kumar Krishna Agrawal, Arna Ghosh, Komal Kumar Teru et al.NeurIPS 2025 · 38 citations
- A distributional simplicity bias in the learning dynamics of transformersRiccardo Rende, Federica Gerace, Alessandro Laio, Sebastian GoldtNeurIPS 2024 · 30 citations
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?Denis Sutter, Julian Minder, Thomas Hofmann, Tiago PimentelNeurIPS 2025 · 30 citations
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and ScaleJames A. Michaelov, Roger P. Levy, Benjamin BergenNeurIPS 2025 · 15 citations
- Sliding Down the Stairs: How Correlated Latent Variables Accelerate Learning with Neural NetworksLorenzo Bardone, Sebastian GoldtICML 2024 · 13 citations
Builds on11
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Frequency Bias in Neural Networks for Input of Non-Uniform DensityRonen Basri, Meirav Galun, Amnon Geifman, David W. Jacobs et al.ICML 2020 · 229 citations
- On the Adequacy of Untuned Warmup for Adaptive OptimizationJerry Ma, Denis YaratsAAAI 2021 · 81 citations
Related papers
- Neural networks trained with SGD learn distributions of increasing complexityMaria Refinetti, Alessandro Ingrosso, Sebastian GoldtICML 2023 · 58 citations
- A Fourier perspective on the learning dynamics of neural networks: from sample complexities to mechanistic insightsFabiola Ricci, Claudia Merger, Sebastian GoldtICML 2026
- Improving Text Generation with Student-Forcing Optimal TransportJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu et al.EMNLP 2020 · 11 citations
- Simplicity Bias of Two-Layer Networks beyond Linearly Separable DataNikita Tsoy, Nikola KonstantinovICML 2024 · 12 citations
- Evading the Simplicity Bias: Training a Diverse Set of Models Discovers Solutions with Superior OOD GeneralizationDamien Teney, Ehsan Abbasnejad, Simon Lucey, Anton van den HengelCVPR 2022 · 32 citations
