Predicting Training Time Without Training
Luca Zancato, Alessandro Achille, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto
Abstract
We tackle the problem of predicting the number of optimization steps that a pretrained deep network needs to converge to a given value of the loss function. To do so, we leverage the fact that the training dynamics of a deep network during fine-tuning are well approximated by those of a linearized model. This allows us to approximate the training loss and accuracy at any point during training by solving a low-dimensional Stochastic Differential Equation (SDE) in function space. Using this result, we are able to predict the time it takes for Stochastic Gradient Descent (SGD) to fine-tune a model to a given loss without having to perform any training. In our experiments, we are able to predict training time of a ResNet within a 20% error margin on a variety of datasets and hyper-parameters, at a 30 to 45-fold reduction in cost compared to actual training. We also discuss how to further reduce the computational and memory cost of our method, and in particular we show that by exploiting the spectral properties of the gradients' matrix it is possible predict training time on a large dataset while processing only a subset of the samples. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22b67b71-948a-48d2-8ece-e52e60a18f77Cited by top-tier papers9
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 272 citations
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 62 citations
- Scaling Neural Tangent Kernels via Sketching and Random FeaturesAmir Zandieh, Insu Han, Haim Avron, Neta Shoham et al.NeurIPS 2021 · 42 citations
- TCT: Convexifying Federated Learning using Bootstrapped Neural Tangent KernelsYaodong Yu, Alexander Wei, Sai Praneeth Karimireddy, Yi Ma et al.NeurIPS 2022 · 38 citations
- Estimating informativeness of samples with Smooth Unique InformationHrayr Harutyunyan, Alessandro Achille, Giovanni Paolini, Orchid Majumder et al.ICLR 2021 · 26 citations
Builds on3
- Rethinking the Hyperparameters for Fine-tuningHao Li, Pratik Chaudhari, Hao Yang, Michael Lam et al.ICLR 2020 · 142 citations
- Gradients as Features for Deep Representation LearningFangzhou Mu, Yingyu Liang, Yin LiICLR 2020 · 46 citations
- Truth or backpropaganda? An empirical investigation of deep learning theoryMicah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller et al.ICLR 2020 · 36 citations
Related papers
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 42 citations
- Locally Regularized Neural Differential Equations: Some Black Boxes were meant to remain closed!Avik Pal, Alan Edelman, Christopher Vincent RackauckasICML 2023 · 4 citations
- LEAD: Exploring Logit Space Evolution for Model SelectionZixuan Hu, Xiaotong Li, Shixiang Tang, Jun Liu et al.CVPR 2024
- Learning Differential Equations that are Easy to SolveJacob Kelly, Jesse Bettencourt, Matthew J. Johnson, David DuvenaudNeurIPS 2020 · 134 citations
