Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
Yanna Ding, Zijie Huang, Xiao Shou, Yihang Guo, Yizhou Sun, Jianxi Gao
Abstract
Learning curve extrapolation predicts neural network performance from early training epochs and has been applied to accelerate AutoML, facilitating hyperparameter tuning and neural architecture search. However, existing methods typically model the evolution of learning curves in isolation, neglecting the impact of neural network (NN) architectures, which influence the loss landscape and learning trajectories. In this work, we explore whether incorporating neural network architecture improves learning curve modeling and how to effectively integrate this architectural information. Motivated by the dynamical system view of optimization, we propose a novel architecture-aware neural differential equation model to forecast learning curves continuously. We empirically demonstrate its ability to capture the general trend of fluctuating learning curves while quantifying uncertainty through variational parameters. Our model outperforms current state-of-the-art learning curve extrapolation methods and pure time-series modeling approaches for both MLP and CNN-based learning curves. Additionally, we explore the applicability of our method in Neural Architecture Search scenarios, such as training configuration ranking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df445636-7f72-449c-b7bf-ac8302eea101Cited by top-tier papers2
- DABO: Difficulty-Aware Bayesian Optimization with Diffusion-Learned PriorsMengyang Li, Pinlong ZhaoCVPR 2026
- Difficulty-Aware Learning Curve ExtrapolationMengyang Li, Pinlong ZhaoAAAI 2026
Builds on12
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 825 citations
- Does Unsupervised Architecture Representation Learning Help Neural Architecture Search?Shen Yan, Yu Zheng, Wei Ao, Xiao Zeng et al.NeurIPS 2020 · 129 citations
- Parameter Prediction for Unseen Deep ArchitecturesBoris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-SorianoNeurIPS 2021 · 111 citations
- Learning Continuous System Dynamics from Irregularly-Sampled Partial ObservationsZijie Huang, Yizhou Sun, Wei WangNeurIPS 2020 · 103 citations
- HOPE: High-order Graph ODE For Modeling Interacting DynamicsXiao Luo, Jingyang Yuan, Zijie Huang, Huiyu Jiang et al.ICML 2023 · 60 citations
Related papers
- Learning to Rank Learning CurvesMartin Wistuba, Tejaswini PedapatiICML 2020 · 31 citations
- How Powerful are Performance Predictors in Neural Architecture Search?Colin White, Arber Zela, Robin Ru, Yang Liu et al.NeurIPS 2021 · 168 citations
- Scaling Laws for Hyperparameter OptimizationArlind Kadra, Maciej Janowski, Martin Wistuba, Josif GrabockaNeurIPS 2023 · 23 citations
- AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the FlyYuchen Jin, Tianyi Zhou, Liangyu Zhao, Yibo Zhu et al.ICLR 2021 · 26 citations
- Neural Architecture and Hyperparameter Selection Through Meta-Learning on Time SeriesErfan Moeini, Christopher Vox, Marie Anastacio, Wadie Skaf et al.AAAI 2026 · 1 citation
