Rethinking Neural Network Learning Rates: A Stackelberg Perspective
Sihan Zeng, Sujay Bhatt, Sumitra Ganesh
Abstract
Neural networks are typically trained with a single learning rate across all layers. While recent empirical evidence suggests that assigning layer-specific learning rates can accelerate training, a principled understanding of the conditions and mechanisms under which non-uniform learning rates are beneficial remains limited. In this work, we investigate non-uniform learning rates through the lens of Stackelberg optimization. Specifically, we demonstrate that training neural networks with a smaller learning rate for the body layers and a larger learning rate for the final layer can be interpreted as a two-time-scale alternating gradient descent algorithm applied to a Stackelberg reformulation of the original objective. We establish finite-time convergence guarantees for the algorithm under broad conditions that accommodate constraint sets and non-smooth activation functions. Beyond convergence, we identify two mechanisms by which non-uniform learning rates can outperform uniform learning rates: (i) we show that certain problem instances induce a Stackelberg objective with stronger optimization structure than the original objective, yielding faster convergence to globally optimal solutions, (ii) our numerical analysis reveals that the Stackelberg objective can exhibit substantially sharper local curvature, especially in early training, which leads to more informative gradients and learning acceleration. Experiments in supervised learning and reinforcement learning support our findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d2618de-9840-4d74-bdc6-2725ebf7fc16Builds on7
- A Fully First-Order Method for Stochastic Bilevel OptimizationJeongyeol Kwon, Dohyun Kwon, Stephen Wright, Robert D. NowakICML 2023 · 123 citations
- Nested Learning: The Illusion of Deep Learning ArchitecturesAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 96 citations
- A Single-timescale Analysis for Stochastic Approximation with Multiple Coupled SequencesHan Shen, Tianyi ChenNeurIPS 2022 · 25 citations
- Leveraging the two-timescale regime to demonstrate convergence of neural networksPierre Marion, Raphaël BerthierNeurIPS 2023 · 19 citations
- A Layer-Wise Natural Gradient Optimizer for Training Deep Neural NetworksXiaolei Liu, Shaoshuai Li, Kaixin Gao, Binfeng WangNeurIPS 2024 · 2 citations
Related papers
- Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural NetworksTianyu Pang, Vignesh Kothapalli, Shenyang Deng, Haohui Wang et al.ICML 2026
- Stackelberg Actor-Critic: Game-Theoretic Reinforcement Learning AlgorithmsLiyuan Zheng, Tanner Fiez, Zane Alumbaugh, Benjamin Chasnov et al.AAAI 2022 · 50 citations
- Convergence of Steepest Descent and Adam under Non-Uniform SmoothnessSharan Vaswani, Yifan Sun, Reza BabanezhadICML 2026 · 1 citation
- Dynamic Game Theoretic Neural OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouICML 2021 · 6 citations
- Independent Policy Gradient Methods for Competitive Reinforcement LearningConstantinos Daskalakis, Dylan J. Foster, Noah GolowichNeurIPS 2020 · 200 citations
