Zero-Shot Performance Prediction for Probabilistic Scaling Laws
Viktoria Schram, Markus Hiller, Daniel Beck, Trevor Cohn
摘要
The prediction of learning curves for Natural Language Processing (NLP) models enables informed decision-making to meet specific performance objectives, while reducing computational overhead and lowering the costs associated with dataset acquisition and curation. In this work, we formulate the prediction task as a multitask learning problem, where each task's data is modelled as being organized within a two-layer hierarchy. To model the shared information and dependencies across tasks and hierarchical levels, we employ latent variable multi-output Gaussian Processes, enabling to account for task correlations and supporting zero-shot prediction of learning curves (LCs). We demonstrate that this approach facilitates the development of probabilistic scaling laws at lower costs. Applying an active learning strategy, LCs can be queried to reduce predictive uncertainty and provide predictions close to ground truth scaling laws. We validate our framework on three small-scale NLP datasets with up to LCs. These are obtained from nanoGPT models, from bilingual translation using mBART and Transformer models, and from multilingual translation using M2M100 models of varying sizes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Active Budget Allocation for Efficient Scaling Law Estimation via Surrogate-Guided PruningViktoria Schram, Markus Hiller, Daniel Beck, Trevor CohnICML 2026
- A Risk Decomposition Framework for Pre-hoc Fine-tuning PredictionYuxiang Luo, Chen Wang, Nan TangICML 2026
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 被引用 171 次
相关 Paper
- Multi Task Learning For Zero Shot Performance Prediction of Multilingual ModelsKabir Ahuja, Shanu Kumar, Sandipan Dandapat, Monojit ChoudhuryACL 2022
- Calibrating Zero-shot Cross-lingual (Un-)structured PredictionsZhengping Jiang, Anqi Liu, Benjamin Van DurmeEMNLP 2022 · 被引用 4 次
- Scaling Multi-Task Bayesian Optimization with Large Language ModelsYimeng Zeng, Natalie Maus, Haydn Thomas Jones, Jeffrey Tao 等ICLR 2026 · 被引用 1 次
- Efficient and Effective Multi-task Grouping via Meta Learning on Task CombinationsXiaozhuang Song, Shun Zheng, Wei Cao, James J. Q. Yu 等NeurIPS 2022 · 被引用 50 次
- Multi-task Causal Learning with Gaussian ProcessesVirginia Aglietti, Theodoros Damoulas, Mauricio A. Álvarez, Javier GonzálezNeurIPS 2020 · 被引用 23 次
