Navigating Scaling Laws: Compute Optimality in Adaptive Model Training
Sotiris Anagnostidis, Gregor Bachmann, Imanol Schlag, Thomas Hofmann
摘要
In recent years, the state-of-the-art in deep learning has been dominated by very large models that have been pre-trained on vast amounts of data. The paradigm is very simple: investing more computational resources (optimally) leads to better performance, and even predictably so; neural scaling laws have been derived that accurately forecast the performance of a network for a desired level of compute. This leads to the notion of a compute-optimal' model, i.e. a model that allocates a given level of compute during training optimally to maximize performance. In this work, we extend the concept of optimality by allowing for an adaptive' model, i.e. a model that can change its shape during training. By doing so, we can design adaptive models that optimally traverse between the underlying scaling laws and outpace their `static' counterparts, leading to a significant reduction in the required compute to reach a given target performance. We show that our approach generalizes across modalities and different shape parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Beyond Next Token Prediction: Patch-Level Training for Large Language ModelsChenze Shao, Fandong Meng, Jie ZhouICLR 2025
- FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less ComputeSotiris Anagnostidis, Gregor Bachmann, Yeongmin Kim, Jonas Kohler 等CVPR 2025
它引用的顶会 Paper33
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
相关 Paper
- Getting ViT in Shape: Scaling Laws for Compute-Optimal Model DesignIbrahim M. Alabdulmohsin, Xiaohua Zhai, Alexander Kolesnikov, Lucas BeyerNeurIPS 2023 · 被引用 122 次
- Scaling View Synthesis TransformersEvan Kim, Hyunwoo Ryu, Thomas W. Mitchel, Vincent SitzmannCVPR 2026 · 被引用 6 次
- Mechanistic Design and Scaling of Hybrid ArchitecturesMichael Poli, Armin W. Thomas, Eric Nguyen, Pragaash Ponnusamy 等ICML 2024 · 被引用 57 次
- Adaptive Data Optimization: Dynamic Sample Selection with Scaling LawsYiding Jiang, Allan Zhou, Zhili Feng, Sadhika Malladi 等ICLR 2025
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and TimeBlake Bordelon, Mary I. Letey, Cengiz PehlevanICLR 2026 · 被引用 14 次
