Budgeted Training: Rethinking Deep Neural Network Training Under Resource Constraints
Mengtian Li, Ersin Yumer, Deva Ramanan
Abstract
In most practical settings and theoretical analyses, one assumes that a model can be trained until convergence. However, the growing complexity of machine learning datasets and models may violate such assumptions. Indeed, current approaches for hyper-parameter tuning and neural architecture search tend to be limited by practical resource constraints. Therefore, we introduce a formal setting for studying training under the non-asymptotic, resource-constrained regime, i.e., budgeted training. We analyze the following problem: "given a dataset, algorithm, and fixed resource budget, what is the best achievable performance?" We focus on the number of optimization iterations as the representative resource. Under such a setting, we show that it is critical to adjust the learning rate schedule according to the given budget. Among budget-aware learning schedules, we find simple linear decay to be both robust and high-performing. We support our claim through extensive experiments with state-of-the-art models on ImageNet (image classification), Kinetics (video classification), MS COCO (object detection and instance segmentation), and Cityscapes (semantic segmentation). We also analyze our results and find that the key to a good schedule is budgeted convergence, a phenomenon whereby the gradient vanishes at the end of each allowed budget. We also revisit existing approaches for fast convergence and show that budget-aware learning schedules readily outperform such approaches under (the practical but under-explored) budgeted training setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1e3a3fa-cb6f-4413-9249-09c6cdc86b57Cited by top-tier papers19
- Train Big, Then Compress: Rethinking Model Size for Efficient Training and Inference of TransformersZhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin et al.ICML 2020 · 184 citations
- Multiplicative Noise and Heavy Tails in Stochastic OptimizationLiam Hodgkinson, Michael W. MahoneyICML 2021 · 90 citations
- FOVEA: Foveated Image Magnification for Autonomous NavigationChittesh Thavamani, Mengtian Li, Nicolas Cebron, Deva RamananICCV 2021 · 45 citations
- EfficientTrain: Exploring Generalized Curriculum Learning for Training Visual BackbonesYulin Wang, Yang Yue, Rui Lu, Tianjiao Liu et al.ICCV 2023 · 39 citations
- Gradient-based Hyperparameter Optimization Over Long HorizonsPaul Micaelli, Amos J. StorkeyNeurIPS 2021 · 23 citations
Related papers
- Stepsize anything: A unified learning rate schedule for budgeted-iteration trainingAnda Tang, Yiming Dong, Yutao Zeng, Xun Zhou et al.NeurIPS 2025 · 1 citation
- How I Learned to Stop Worrying and Love RetrainingMax Zimmer, Christoph Spiegel, Sebastian PokuttaICLR 2023
- Learning Rate Annealing Improves Tuning Robustness in Stochastic OptimizationAmit Attia, Tomer KorenICML 2026
- Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained ComputationWenxuan Zhang, Youssef Mohamed, Bernard Ghanem, Philip Torr et al.ICLR 2024 · 6 citations
- The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model TrainingFabian Schaipp, Alexander Hägele, Adrien B. Taylor, Umut Simsekli et al.ICML 2025
