Budgeted Training: Rethinking Deep Neural Network Training Under Resource Constraints
Mengtian Li, Ersin Yumer, Deva Ramanan
摘要
In most practical settings and theoretical analyses, one assumes that a model can be trained until convergence. However, the growing complexity of machine learning datasets and models may violate such assumptions. Indeed, current approaches for hyper-parameter tuning and neural architecture search tend to be limited by practical resource constraints. Therefore, we introduce a formal setting for studying training under the non-asymptotic, resource-constrained regime, i.e., budgeted training. We analyze the following problem: "given a dataset, algorithm, and fixed resource budget, what is the best achievable performance?" We focus on the number of optimization iterations as the representative resource. Under such a setting, we show that it is critical to adjust the learning rate schedule according to the given budget. Among budget-aware learning schedules, we find simple linear decay to be both robust and high-performing. We support our claim through extensive experiments with state-of-the-art models on ImageNet (image classification), Kinetics (video classification), MS COCO (object detection and instance segmentation), and Cityscapes (semantic segmentation). We also analyze our results and find that the key to a good schedule is budgeted convergence, a phenomenon whereby the gradient vanishes at the end of each allowed budget. We also revisit existing approaches for fast convergence and show that budget-aware learning schedules readily outperform such approaches under (the practical but under-explored) budgeted training setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Train Big, Then Compress: Rethinking Model Size for Efficient Training and Inference of TransformersZhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin 等ICML 2020 · 被引用 184 次
- Multiplicative Noise and Heavy Tails in Stochastic OptimizationLiam Hodgkinson, Michael W. MahoneyICML 2021 · 被引用 90 次
- FOVEA: Foveated Image Magnification for Autonomous NavigationChittesh Thavamani, Mengtian Li, Nicolas Cebron, Deva RamananICCV 2021 · 被引用 45 次
- EfficientTrain: Exploring Generalized Curriculum Learning for Training Visual BackbonesYulin Wang, Yang Yue, Rui Lu, Tianjiao Liu 等ICCV 2023 · 被引用 39 次
- Gradient-based Hyperparameter Optimization Over Long HorizonsPaul Micaelli, Amos J. StorkeyNeurIPS 2021 · 被引用 23 次
相关 Paper
- Stepsize anything: A unified learning rate schedule for budgeted-iteration trainingAnda Tang, Yiming Dong, Yutao Zeng, Xun Zhou 等NeurIPS 2025 · 被引用 1 次
- How I Learned to Stop Worrying and Love RetrainingMax Zimmer, Christoph Spiegel, Sebastian PokuttaICLR 2023
- Learning Rate Annealing Improves Tuning Robustness in Stochastic OptimizationAmit Attia, Tomer KorenICML 2026
- Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained ComputationWenxuan Zhang, Youssef Mohamed, Bernard Ghanem, Philip Torr 等ICLR 2024 · 被引用 6 次
- The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model TrainingFabian Schaipp, Alexander Hägele, Adrien B. Taylor, Umut Simsekli 等ICML 2025
