Escaping the Gravitational Pull of Softmax
Jincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li, Csaba Szepesvári, Dale Schuurmans
Abstract
The softmax is the standard transformation used in machine learning to map realvalued vectors to categorical distributions. Unfortunately, this transform poses serious drawbacks for gradient descent (ascent) optimization. We reveal this difficulty by establishing two negative results: (1) optimizing any expectation with respect to the softmax must exhibit sensitivity to parameter initialization ("softmax gravity well"), and (2) optimizing log-probabilities under the softmax must exhibit slow convergence ("softmax damping"). Both findings are based on an analysis of convergence rates using the Non-uniform Łojasiewicz (NŁ) inequalities. To circumvent these shortcomings we investigate an alternative transformation, the escort mapping, that demonstrates better optimization properties. The disadvantages of the softmax and the effectiveness of the escort transformation are further explained using the concept of NŁ coefficient. In addition to proving bounds on convergence rates to firmly establish these results, we also provide experimental evidence for the superiority of the escort transformation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86c8ef5c-b5f5-4ff5-8493-0d2728364270Cited by top-tier papers19
- What Makes a Reward Model a Good Teacher? An Optimization PerspectiveNoam Razin, Zixuan Wang, Hubert Strauss, Stanley Wei et al.NeurIPS 2025 · 73 citations
- Controlling Graph Dynamics with Reinforcement Learning and Graph Neural NetworksEli A. Meirom, Haggai Maron, Shie Mannor, Gal ChechikICML 2021 · 56 citations
- Leveraging Non-uniformity in First-order Non-convex OptimizationJincheng Mei, Yue Gao, Bo Dai, Csaba Szepesvári et al.ICML 2021 · 55 citations
- The Role of Baselines in Policy Gradient OptimizationJincheng Mei, Wesley Chung, Valentin Thomas, Bo Dai et al.NeurIPS 2022 · 34 citations
- Vanishing Gradients in Reinforcement Finetuning of Language ModelsNoam Razin, Hattie Zhou, Omid Saremi, Vimal Thilak et al.ICLR 2024 · 27 citations
Builds on1
Related papers
- Enhancing Classifier Conservativeness and Robustness by PolynomialityZiqi Wang, Marco LoogCVPR 2022 · 2 citations
- Gradient-Normalized Smoothness for Optimization with Approximate HessiansAndrei Semenov, Martin Jaggi, Nikita DoikovICLR 2026 · 8 citations
- Armijo Line-search Can Make (Stochastic) Gradient Descent Provably FasterSharan Vaswani, Reza Babanezhad HarikandehICML 2025
- MultiMax: Sparse and Multi-Modal Attention LearningYuxuan Zhou, Mario Fritz, Margret KeuperICML 2024 · 4 citations
- Extreme Classification via Adversarial Softmax ApproximationRobert Bamler, Stephan MandtICLR 2020 · 25 citations
