Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
Yanna Ding, Zijie Huang, Xiao Shou, Yihang Guo, Yizhou Sun, Jianxi Gao
摘要
Learning curve extrapolation predicts neural network performance from early training epochs and has been applied to accelerate AutoML, facilitating hyperparameter tuning and neural architecture search. However, existing methods typically model the evolution of learning curves in isolation, neglecting the impact of neural network (NN) architectures, which influence the loss landscape and learning trajectories. In this work, we explore whether incorporating neural network architecture improves learning curve modeling and how to effectively integrate this architectural information. Motivated by the dynamical system view of optimization, we propose a novel architecture-aware neural differential equation model to forecast learning curves continuously. We empirically demonstrate its ability to capture the general trend of fluctuating learning curves while quantifying uncertainty through variational parameters. Our model outperforms current state-of-the-art learning curve extrapolation methods and pure time-series modeling approaches for both MLP and CNN-based learning curves. Additionally, we explore the applicability of our method in Neural Architecture Search scenarios, such as training configuration ranking.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- DABO: Difficulty-Aware Bayesian Optimization with Diffusion-Learned PriorsMengyang Li, Pinlong ZhaoCVPR 2026
- Difficulty-Aware Learning Curve ExtrapolationMengyang Li, Pinlong ZhaoAAAI 2026
它引用的顶会 Paper12
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 被引用 825 次
- Does Unsupervised Architecture Representation Learning Help Neural Architecture Search?Shen Yan, Yu Zheng, Wei Ao, Xiao Zeng 等NeurIPS 2020 · 被引用 129 次
- Parameter Prediction for Unseen Deep ArchitecturesBoris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-SorianoNeurIPS 2021 · 被引用 111 次
- Learning Continuous System Dynamics from Irregularly-Sampled Partial ObservationsZijie Huang, Yizhou Sun, Wei WangNeurIPS 2020 · 被引用 103 次
- HOPE: High-order Graph ODE For Modeling Interacting DynamicsXiao Luo, Jingyang Yuan, Zijie Huang, Huiyu Jiang 等ICML 2023 · 被引用 60 次
相关 Paper
- Learning to Rank Learning CurvesMartin Wistuba, Tejaswini PedapatiICML 2020 · 被引用 31 次
- How Powerful are Performance Predictors in Neural Architecture Search?Colin White, Arber Zela, Robin Ru, Yang Liu 等NeurIPS 2021 · 被引用 168 次
- Scaling Laws for Hyperparameter OptimizationArlind Kadra, Maciej Janowski, Martin Wistuba, Josif GrabockaNeurIPS 2023 · 被引用 23 次
- AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the FlyYuchen Jin, Tianyi Zhou, Liangyu Zhao, Yibo Zhu 等ICLR 2021 · 被引用 26 次
- Neural Architecture and Hyperparameter Selection Through Meta-Learning on Time SeriesErfan Moeini, Christopher Vox, Marie Anastacio, Wadie Skaf 等AAAI 2026 · 被引用 1 次
