Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted Networks
Steven Adriaensen, Herilalaina Rakotoarison, Samuel Müller, Frank Hutter
摘要
Learning curve extrapolation aims to predict model performance in later epochs of training, based on the performance in earlier epochs. In this work, we argue that, while the inherent uncertainty in the extrapolation of learning curves warrants a Bayesian approach, existing methods are (i) overly restrictive, and/or (ii) computationally expensive. We describe the first application of prior-data fitted neural networks (PFNs) in this context. A PFN is a transformer, pre-trained on data generated from a prior, to perform approximate Bayesian inference in a single forward pass. We propose LC-PFN, a PFN trained to extrapolate artificial right-censored learning curves generated from a parametric prior proposed in prior art using MCMC. We demonstrate that LC-PFN can approximate the posterior predictive distribution over learning curves more accurately than MCMC, while being over 10 000 times faster. We also show that the same LC-PFN achieves competitive performance extrapolating a total of 20 000 real learning curves from four learning curve benchmarks (LCBench, NAS-Bench-201, Taskset, and PD1) that stem from training a wide range of model architectures (MLPs, CNNs, RNNs, and Transformers) on 53 different datasets with varying input modalities (tabular, image, text, and protein data). Finally, we investigate its potential in the context of model selection and find that a simple LC-PFN based predictive early stopping criterion obtains 2 -6× speed-ups on 45 of these datasets, at virtually no overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- ForecastPFN: Synthetically-Trained Zero-Shot ForecastingSamuel Dooley, Gurnoor Singh Khurana, Chirag Mohapatra, Siddartha V. Naidu 等NeurIPS 2023 · 被引用 142 次
- TuneTables: Context Optimization for Scalable Prior-Data Fitted NetworksBenjamin Feuer, Robin Schirrmeister, Valeriia Cherepanova, Chinmay Hegde 等NeurIPS 2024 · 被引用 57 次
- In-Context Freeze-Thaw Bayesian Optimization for Hyperparameter OptimizationHerilalaina Rakotoarison, Steven Adriaensen, Neeratyoy Mallik, Samir Garibov 等ICML 2024 · 被引用 28 次
- Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential EquationYanna Ding, Zijie Huang, Xiao Shou, Yihang Guo 等AAAI 2025 · 被引用 4 次
- Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter TuningDong Bok Lee, Aoxuan Silvia Zhang, Byungjoo Kim, Junhyeon Park 等NeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper6
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 被引用 825 次
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka 等ICLR 2022 · 被引用 287 次
- ForecastPFN: Synthetically-Trained Zero-Shot ForecastingSamuel Dooley, Gurnoor Singh Khurana, Chirag Mohapatra, Siddartha V. Naidu 等NeurIPS 2023 · 被引用 142 次
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 被引用 96 次
- PFNs4BO: In-Context Learning for Bayesian OptimizationSamuel Müller, Matthias Feurer, Noah Hollmann, Frank HutterICML 2023 · 被引用 71 次
相关 Paper
- Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted NetworksDongwoo Lee, Dong Bok Lee, Steven Adriaensen, Juho Lee 等ICML 2025
- Statistical Foundations of Prior-Data Fitted NetworksThomas NaglerICML 2023 · 被引用 51 次
- TimePFN: Effective Multivariate Time Series Forecasting with Synthetic DataEge Onur Taga, Muhammed Emrullah Ildiz, Samet OymakAAAI 2025 · 被引用 27 次
- Difficulty-Aware Learning Curve ExtrapolationMengyang Li, Pinlong ZhaoAAAI 2026
- Transformers Can Learn Posterior Predictive Distributions In-ContextGyeonghun Kang, Changwoo Lee, Xiang ChengICML 2026
