Training Neural Networks for and by Interpolation
Leonard Berrada, Andrew Zisserman, M. Pawan Kumar
摘要
The majority of modern deep learning models are able to interpolate the data: the empirical loss can be driven near zero on all samples simultaneously. In this work, we explicitly exploit this interpolation property for the design of a new optimization algorithm for deep learning. Specifically, we use it to compute an adaptive learning-rate given a stochastic gradient direction. This results in the Adaptive Learning-rates for Interpolation with Gradients (ALI-G) algorithm. ALI-G retains the advantages of SGD, which are low computational cost and provable convergence in the convex setting. But unlike SGD, the learning-rate of ALI-G can be computed inexpensively in closed-form and does not require a manual schedule. We provide a detailed analysis of ALI-G in the stochastic convex setting with explicit convergence rates. In order to obtain good empirical performance in deep learning, we extend the algorithm to use a maximal learning-rate, which gives a single hyper-parameter to tune. We show that employing such a maximal learning-rate has an intuitive proximal interpretation and preserves all convergence guarantees. We provide experiments on a variety of architectures and tasks: (i) learning a differentiable neural computer; (ii) training a wide residual network on the SVHN data set; (iii) training a Bi-LSTM on the SNLI data set; and (iv) training wide residual networks and densely connected networks on the CIFAR data sets. We empirically show that ALI-G outperforms adaptive gradient methods such as Adam, and provides comparable performance with SGD, although SGD benefits from manual learning rate schedules. We release PyTorch and Tensorflow implementations of ALI-G as standalone optimizers that can be used as a drop-in replacement in existing code (code available at this https URL ).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 被引用 195 次
- DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size ScheduleMaor Ivgi, Oliver Hinder, Yair CarmonICML 2023 · 被引用 98 次
- Dynamics of SGD with Stochastic Polyak Stepsizes: Truly Adaptive Variants and Convergence to Exact SolutionAntonio Orvieto, Simon Lacoste-Julien, Nicolas LoizouNeurIPS 2022 · 被引用 57 次
- Adaptive SGD with Polyak stepsize and Line-search: Robust Convergence and Variance ReductionXiaowen Jiang, Sebastian U. StichNeurIPS 2023 · 被引用 40 次
- Generalized Polyak Step Size for First Order Optimization with MomentumXiaoyu Wang, Mikael Johansson, Tong ZhangICML 2023 · 被引用 32 次
它引用的顶会 Paper3
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Accelerating SGD with momentum for over-parameterized learningChaoyue Liu, Mikhail BelkinICLR 2020 · 被引用 93 次
- Small Steps and Giant Leaps: Minimal Newton Solvers for Deep LearningJoão F. Henriques, Sébastien Ehrhardt, Samuel Albanie, Andrea VedaldiICCV 2019 · 被引用 23 次
相关 Paper
- BiSLS/SPS: Auto-tune Step Sizes for Stable Bi-level OptimizationChen Fan, Gaspard Choné-Ducasse, Mark Schmidt, Christos ThrampoulidisNeurIPS 2023 · 被引用 6 次
- SUPER-ADAM: Faster and Universal Framework of Adaptive GradientsFeihu Huang, Junyi Li, Heng HuangNeurIPS 2021 · 被引用 55 次
- The Implicit Bias for Adaptive Optimization Algorithms on Homogeneous Neural NetworksBohan Wang, Qi Meng, Wei Chen, Tie-Yan LiuICML 2021 · 被引用 45 次
- Extrapolation for Large-batch Training in Deep LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiICML 2020 · 被引用 43 次
- Provable Adaptivity of Adam under Non-uniform SmoothnessBohan Wang, Yushun Zhang, Huishuai Zhang, Qi Meng 等KDD 2024 · 被引用 4 次
