Data Efficient Neural Scaling Law via Model Reusing
Peihao Wang, Rameswar Panda, Zhangyang Wang
摘要
The number of parameters in large transformers has been observed to grow exponentially. Despite notable performance improvements, concerns have been raised that such a growing model size will run out of data in the near future. As manifested in the neural scaling law, modern learning backbones are not data-efficient. To maintain the utility of the model capacity, training data should be increased proportionally. In this paper, we study the neural scaling law under the previously overlooked data scarcity regime, focusing on the more challenging situation where we need to train a gigantic model with a disproportionately limited supply of available training data. We find that the existing power laws underestimate the data inefficiency of large transformers. Their performance will drop significantly if the training set is insufficient. Fortunately, we discover another blessing -such a data-inefficient scaling law can be restored through a model reusing approach that warm-starts the training of a large model by initializing it using smaller models. Our empirical study shows that model reusing can effectively reproduce the power law under the data scarcity regime. When progressively applying model reusing to expand the model size, we also observe consistent performance improvement in large transformers. We release our code at: https://github.com/VITA-Group/ Data-Efficient-Scaling .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Diffusion Model as Representation LearnerXingyi Yang, Xinchao WangICCV 2023 · 被引用 100 次
- Flextron: Many-in-One Flexible Large Language ModelRuisi Cai, Saurav Muralidharan, Greg Heinrich, Hongxu Yin 等ICML 2024 · 被引用 38 次
- Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear RegressionTingkai Yan, Haodong Wen, Binghui Li, Kairong Luo 等ICLR 2026 · 被引用 12 次
- Initializing Variable-sized Vision Transformers from Learngene with Learnable TransformationShiyu Xia, Yuankun Zu, Xu Yang, Xin GengNeurIPS 2024 · 被引用 9 次
- Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMsHouyi Li, Wenzhen Zheng, Qiufeng Wang, Zhenyu Ding 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,279 次
相关 Paper
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- Scaling Laws for Upcycling Mixture-of-Experts Language ModelsSeng Pei Liew, Takuya Kato, Sho TakaseICML 2025
- Learning to Grow Pretrained Models for Efficient Transformer TrainingPeihao Wang, Rameswar Panda, Lucas Torroba Hennigen, Philip Greengard 等ICLR 2023 · 被引用 13 次
- Scaling Laws for Sparsely-Connected Foundation ModelsElias Frantar, Carlos Riquelme Ruiz, Neil Houlsby, Dan Alistarh 等ICLR 2024 · 被引用 48 次
- Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training StrategiesZhengyu Chen, Siqi Wang, Teng Xiao, Yudong Wang 等ACL 2025
