Gradient descent with generalized Newton's method
Zhiqi Bu, Shiyun Xu
Abstract
We propose the generalized Newton's method (GeN) -a Hessian-informed approach that applies to any optimizer such as SGD and Adam, and covers the Newton-Raphson method as a sub-case. Our method automatically and dynamically selects the learning rate that accelerates the convergence, without the intensive tuning of the learning rate scheduler. In practice, our method is easily implementable, since it only requires additional forward passes with almost zero computational overhead (in terms of training time and memory cost), if the overhead is amortized over many iterations. We present extensive experiments on language and vision tasks (e.g. GPT and ResNet) to showcase that GeN optimizers match the state-of-the-art performance, which was achieved with carefully tuned learning rate schedulers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 289ceb5c-17f7-4a35-9155-4a66af04e296Cited by top-tier papers3
- Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning RateZhiqi Bu, Shiyun Xu, Jialin MaoICLR 2026 · 4 citations
- Scaling depth capacity via zero/one-layer model expansionZhiqi BuICML 2026 · 1 citation
- Towards hyperparameter-free optimization with differential privacyRuixuan Liu, Zhiqi BuICLR 2025
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- Domain-Independent Dominance of Adaptive MethodsPedro Savarese, David McAllester, Sudarshan Babu, Michael MaireCVPR 2021
- Variational Learning is Effective for Large Deep NetworksYuesong Shen, Nico Daheim, Bai Cong, Peter Nickl et al.ICML 2024 · 53 citations
- MADA: Meta-Adaptive Optimizers Through Hyper-Gradient DescentKaan Ozkara, Can Karakus, Parameswaran Raman, Mingyi Hong et al.ICML 2024 · 7 citations
- Learning to Schedule Learning rate with Graph Neural NetworksYuanhao Xiong, Li-Cheng Lan, Xiangning Chen, Ruochen Wang et al.ICLR 2022 · 18 citations
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 117 citations
