Learning curves theory for hierarchically compositional data with power-law distributed features
Francesco Cagnetta, Hyunmo Kang, Matthieu Wyart
摘要
Recent theories suggest that Neural Scaling Laws arise whenever the task is linearly decomposed into power-law distributed units. Alternatively, scaling laws also emerge when data exhibit a hierarchically compositional structure, as is thought to occur in language and images. To unify these views, we consider classification and next-token prediction tasks based on probabilistic contextfree grammars-probabilistic models that generate data via a hierarchy of production rules. For classification, we show that having power-law distributed production rules results in a power-law learning curve with an exponent depending on the rules' distribution and a large multiplicative constant that depends on the hierarchical structure. By contrast, for next-token prediction, the distribution of production rules controls the local details of the learning curve, but not the exponent describing the large-scale behaviour.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Spectral Reach: Understanding Neural Scaling as Progress into the Spectral TailKonstantin Nikolaou, Jonas Scheunemann, Sven Krippendorf, Samuel Tovey 等ICML 2026
- Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-trainingLei Liu, Hao Zhu, Xiaoyan Yang, Yue Shen 等ACL 2026
- Diffusion Models Preferentially Memorize Prototypical Examples or: Why Does My Diffusion Model Love Slop?Marta Aparicio Rodriguez, Anastasia Borovykh, Grigorios A Pavliotis, Daniel KorchinskiICML 2026
它引用的顶会 Paper16
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 被引用 245 次
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 被引用 179 次
- A Dynamical Model of Neural Scaling LawsBlake Bordelon, Alexander B. Atanasov, Cengiz PehlevanICML 2024 · 被引用 84 次
- Transformers Represent Belief State Geometry in their Residual StreamAdam S. Shai, Lucas Teixeira, Alexander Gietelink Oldenziel, Sarah Marzen 等NeurIPS 2024 · 被引用 83 次
相关 Paper
- Towards a theory of how the structure of language is acquired by deep neural networksFrancesco Cagnetta, Matthieu WyartNeurIPS 2024 · 被引用 33 次
- Deriving Neural Scaling Laws from the Statistics of Natural LanguageFrancesco Cagnetta, Allan Raventos, Surya Ganguli, Matthieu WyartICML 2026
- On the origin of neural scaling laws: from random graphs to natural languageMaissam Barkeshli, Alberto Alfarano, Andrey GromovICML 2026
- How Compositional Generalization and Creativity Improve as Diffusion Models are TrainedAlessandro Favero, Antonio Sclocchi, Francesco Cagnetta, Pascal Frossard 等ICML 2025
- Deep networks learn to parse uniform-depth context-free languages from local statisticsJack T. Parley, Francesco Cagnetta, Matthieu WyartICML 2026 · 被引用 4 次
