Gradient descent with generalized Newton's method
Zhiqi Bu, Shiyun Xu
摘要
We propose the generalized Newton's method (GeN) -a Hessian-informed approach that applies to any optimizer such as SGD and Adam, and covers the Newton-Raphson method as a sub-case. Our method automatically and dynamically selects the learning rate that accelerates the convergence, without the intensive tuning of the learning rate scheduler. In practice, our method is easily implementable, since it only requires additional forward passes with almost zero computational overhead (in terms of training time and memory cost), if the overhead is amortized over many iterations. We present extensive experiments on language and vision tasks (e.g. GPT and ResNet) to showcase that GeN optimizers match the state-of-the-art performance, which was achieved with carefully tuned learning rate schedulers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning RateZhiqi Bu, Shiyun Xu, Jialin MaoICLR 2026 · 被引用 4 次
- Scaling depth capacity via zero/one-layer model expansionZhiqi BuICML 2026 · 被引用 1 次
- Towards hyperparameter-free optimization with differential privacyRuixuan Liu, Zhiqi BuICLR 2025
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
相关 Paper
- Domain-Independent Dominance of Adaptive MethodsPedro Savarese, David McAllester, Sudarshan Babu, Michael MaireCVPR 2021
- Variational Learning is Effective for Large Deep NetworksYuesong Shen, Nico Daheim, Bai Cong, Peter Nickl 等ICML 2024 · 被引用 53 次
- MADA: Meta-Adaptive Optimizers Through Hyper-Gradient DescentKaan Ozkara, Can Karakus, Parameswaran Raman, Mingyi Hong 等ICML 2024 · 被引用 7 次
- Learning to Schedule Learning rate with Graph Neural NetworksYuanhao Xiong, Li-Cheng Lan, Xiangning Chen, Ruochen Wang 等ICLR 2022 · 被引用 18 次
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 被引用 117 次
