Predicting Training Time Without Training
Luca Zancato, Alessandro Achille, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto
摘要
We tackle the problem of predicting the number of optimization steps that a pretrained deep network needs to converge to a given value of the loss function. To do so, we leverage the fact that the training dynamics of a deep network during fine-tuning are well approximated by those of a linearized model. This allows us to approximate the training loss and accuracy at any point during training by solving a low-dimensional Stochastic Differential Equation (SDE) in function space. Using this result, we are able to predict the time it takes for Stochastic Gradient Descent (SGD) to fine-tune a model to a given loss without having to perform any training. In our experiments, we are able to predict training time of a ResNet within a 20% error margin on a variety of datasets and hyper-parameters, at a 30 to 45-fold reduction in cost compared to actual training. We also discuss how to further reduce the computational and memory cost of our method, and in particular we show that by exploiting the spectral properties of the gradients' matrix it is possible predict training time on a large dataset while processing only a subset of the samples. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 被引用 272 次
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 被引用 62 次
- Scaling Neural Tangent Kernels via Sketching and Random FeaturesAmir Zandieh, Insu Han, Haim Avron, Neta Shoham 等NeurIPS 2021 · 被引用 42 次
- TCT: Convexifying Federated Learning using Bootstrapped Neural Tangent KernelsYaodong Yu, Alexander Wei, Sai Praneeth Karimireddy, Yi Ma 等NeurIPS 2022 · 被引用 38 次
- Estimating informativeness of samples with Smooth Unique InformationHrayr Harutyunyan, Alessandro Achille, Giovanni Paolini, Orchid Majumder 等ICLR 2021 · 被引用 26 次
它引用的顶会 Paper3
- Rethinking the Hyperparameters for Fine-tuningHao Li, Pratik Chaudhari, Hao Yang, Michael Lam 等ICLR 2020 · 被引用 142 次
- Gradients as Features for Deep Representation LearningFangzhou Mu, Yingyu Liang, Yin LiICLR 2020 · 被引用 46 次
- Truth or backpropaganda? An empirical investigation of deep learning theoryMicah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller 等ICLR 2020 · 被引用 36 次
相关 Paper
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 被引用 42 次
- Locally Regularized Neural Differential Equations: Some Black Boxes were meant to remain closed!Avik Pal, Alan Edelman, Christopher Vincent RackauckasICML 2023 · 被引用 4 次
- LEAD: Exploring Logit Space Evolution for Model SelectionZixuan Hu, Xiaotong Li, Shixiang Tang, Jun Liu 等CVPR 2024
- Learning Differential Equations that are Easy to SolveJacob Kelly, Jesse Bettencourt, Matthew J. Johnson, David DuvenaudNeurIPS 2020 · 被引用 134 次
