Difficulty-Aware Learning Curve Extrapolation
Mengyang Li, Pinlong Zhao
Abstract
Learning Curve Extrapolation (LCE) is a critical technique for accelerating automated machine learning by terminating unpromising training runs early. Recent state-of-the-art methods have improved predictive accuracy by incorporating contextual information, such as neural network architecture. However, these approaches, whether context-agnostic or architecture-aware, still operate under the implicit assumption of a uniform task landscape. They overlook a pivotal, complementary factor: the intrinsic difficulty of the learning task itself. This oversight leads to significant performance degradation, especially for tasks whose learning dynamics diverge from the model's priors. In this work, we argue that task difficulty is a crucial yet neglected dimension for robust LCE. We introduce Difficulty-Aware Learning Curve Extrapolation (DA-LCE), which explicitly conditions its predictions on task complexity. Our core contributions are threefold: (1) We propose a transparent, rule-based method to quantify task difficulty from early learning curve dynamics, eliminating the need for external meta-features. (2) We design a novel data generation pipeline using conditional diffusion models to create high-fidelity, difficulty-conditioned synthetic training data. (3) We introduce a Transformer-based predictor that leverages difficulty information to achieve superior accuracy across diverse benchmarks. Extensive experiments demonstrate that our approach significantly outperforms both difficulty-agnostic and architecture-aware baselines, with task difficulty emerging as a powerful conditioning signal whose impact matches or exceeds that of model architecture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Dual-Difficulty Curriculum Learning for Direct Preference OptimizationMengyang Li, Haozhan Geng, Zhong Zhang, Shuang LiuKDD 2026 · 5 citations
- What Do LLMs Learn First? Asymmetric Learning Dynamics of Input Complexity and Output Ambiguity in Preference AlignmentMengyang Li, Jingwen Wang, Pinlong ZhaoACL 2026
- Layer-wise Gradient Disentanglement: Decoupling Semantics and Preferences in Direct Preference OptimizationMengyang Li, Shuang Liu, Zhong ZhangICML 2026
Builds on5
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Surrogate NAS Benchmarks: Going Beyond the Limited Search Spaces of Tabular NAS BenchmarksArber Zela, Julien Niklas Siems, Lucas Zimmer, Jovita Lukasik et al.ICLR 2022 · 100 citations
- Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted NetworksSteven Adriaensen, Herilalaina Rakotoarison, Samuel Müller, Frank HutterNeurIPS 2023 · 55 citations
- Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential EquationYanna Ding, Zijie Huang, Xiao Shou, Yihang Guo et al.AAAI 2025 · 4 citations
Related papers
- Learning to Anticipate Future with Dynamic Context RemovalXinyu Xu, Yong-Lu Li, Cewu LuCVPR 2022 · 18 citations
- DABO: Difficulty-Aware Bayesian Optimization with Diffusion-Learned PriorsMengyang Li, Pinlong ZhaoCVPR 2026
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 171 citations
- Uncertainty-Aware Curriculum Learning for Neural Machine TranslationYikai Zhou, Baosong Yang, Derek F. Wong, Yu Wan et al.ACL 2020 · 78 citations
- Generative Adaptation of Dynamics to Environmental Shifts via Weight-space DiffusionRuikun Li, Huandong Wang, Jingtao Ding, Yuan Yuan et al.ICML 2026 · 4 citations
