Difficulty-Aware Learning Curve Extrapolation
Mengyang Li, Pinlong Zhao
摘要
Learning Curve Extrapolation (LCE) is a critical technique for accelerating automated machine learning by terminating unpromising training runs early. Recent state-of-the-art methods have improved predictive accuracy by incorporating contextual information, such as neural network architecture. However, these approaches, whether context-agnostic or architecture-aware, still operate under the implicit assumption of a uniform task landscape. They overlook a pivotal, complementary factor: the intrinsic difficulty of the learning task itself. This oversight leads to significant performance degradation, especially for tasks whose learning dynamics diverge from the model's priors. In this work, we argue that task difficulty is a crucial yet neglected dimension for robust LCE. We introduce Difficulty-Aware Learning Curve Extrapolation (DA-LCE), which explicitly conditions its predictions on task complexity. Our core contributions are threefold: (1) We propose a transparent, rule-based method to quantify task difficulty from early learning curve dynamics, eliminating the need for external meta-features. (2) We design a novel data generation pipeline using conditional diffusion models to create high-fidelity, difficulty-conditioned synthetic training data. (3) We introduce a Transformer-based predictor that leverages difficulty information to achieve superior accuracy across diverse benchmarks. Extensive experiments demonstrate that our approach significantly outperforms both difficulty-agnostic and architecture-aware baselines, with task difficulty emerging as a powerful conditioning signal whose impact matches or exceeds that of model architecture.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Dual-Difficulty Curriculum Learning for Direct Preference OptimizationMengyang Li, Haozhan Geng, Zhong Zhang, Shuang LiuKDD 2026 · 被引用 5 次
- What Do LLMs Learn First? Asymmetric Learning Dynamics of Input Complexity and Output Ambiguity in Preference AlignmentMengyang Li, Jingwen Wang, Pinlong ZhaoACL 2026
- Layer-wise Gradient Disentanglement: Decoupling Semantics and Preferences in Direct Preference OptimizationMengyang Li, Shuang Liu, Zhong ZhangICML 2026
它引用的顶会 Paper5
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Surrogate NAS Benchmarks: Going Beyond the Limited Search Spaces of Tabular NAS BenchmarksArber Zela, Julien Niklas Siems, Lucas Zimmer, Jovita Lukasik 等ICLR 2022 · 被引用 100 次
- Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted NetworksSteven Adriaensen, Herilalaina Rakotoarison, Samuel Müller, Frank HutterNeurIPS 2023 · 被引用 55 次
- Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential EquationYanna Ding, Zijie Huang, Xiao Shou, Yihang Guo 等AAAI 2025 · 被引用 4 次
相关 Paper
- Learning to Anticipate Future with Dynamic Context RemovalXinyu Xu, Yong-Lu Li, Cewu LuCVPR 2022 · 被引用 18 次
- DABO: Difficulty-Aware Bayesian Optimization with Diffusion-Learned PriorsMengyang Li, Pinlong ZhaoCVPR 2026
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 被引用 171 次
- Uncertainty-Aware Curriculum Learning for Neural Machine TranslationYikai Zhou, Baosong Yang, Derek F. Wong, Yu Wan 等ACL 2020 · 被引用 78 次
- Generative Adaptation of Dynamics to Environmental Shifts via Weight-space DiffusionRuikun Li, Huandong Wang, Jingtao Ding, Yuan Yuan 等ICML 2026 · 被引用 4 次
