Escaping the Gravitational Pull of Softmax
Jincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li, Csaba Szepesvári, Dale Schuurmans
摘要
The softmax is the standard transformation used in machine learning to map realvalued vectors to categorical distributions. Unfortunately, this transform poses serious drawbacks for gradient descent (ascent) optimization. We reveal this difficulty by establishing two negative results: (1) optimizing any expectation with respect to the softmax must exhibit sensitivity to parameter initialization ("softmax gravity well"), and (2) optimizing log-probabilities under the softmax must exhibit slow convergence ("softmax damping"). Both findings are based on an analysis of convergence rates using the Non-uniform Łojasiewicz (NŁ) inequalities. To circumvent these shortcomings we investigate an alternative transformation, the escort mapping, that demonstrates better optimization properties. The disadvantages of the softmax and the effectiveness of the escort transformation are further explained using the concept of NŁ coefficient. In addition to proving bounds on convergence rates to firmly establish these results, we also provide experimental evidence for the superiority of the escort transformation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- What Makes a Reward Model a Good Teacher? An Optimization PerspectiveNoam Razin, Zixuan Wang, Hubert Strauss, Stanley Wei 等NeurIPS 2025 · 被引用 73 次
- Controlling Graph Dynamics with Reinforcement Learning and Graph Neural NetworksEli A. Meirom, Haggai Maron, Shie Mannor, Gal ChechikICML 2021 · 被引用 56 次
- Leveraging Non-uniformity in First-order Non-convex OptimizationJincheng Mei, Yue Gao, Bo Dai, Csaba Szepesvári 等ICML 2021 · 被引用 55 次
- The Role of Baselines in Policy Gradient OptimizationJincheng Mei, Wesley Chung, Valentin Thomas, Bo Dai 等NeurIPS 2022 · 被引用 34 次
- Vanishing Gradients in Reinforcement Finetuning of Language ModelsNoam Razin, Hattie Zhou, Omid Saremi, Vimal Thilak 等ICLR 2024 · 被引用 27 次
它引用的顶会 Paper1
相关 Paper
- Enhancing Classifier Conservativeness and Robustness by PolynomialityZiqi Wang, Marco LoogCVPR 2022 · 被引用 2 次
- Gradient-Normalized Smoothness for Optimization with Approximate HessiansAndrei Semenov, Martin Jaggi, Nikita DoikovICLR 2026 · 被引用 8 次
- Armijo Line-search Can Make (Stochastic) Gradient Descent Provably FasterSharan Vaswani, Reza Babanezhad HarikandehICML 2025
- MultiMax: Sparse and Multi-Modal Attention LearningYuxuan Zhou, Mario Fritz, Margret KeuperICML 2024 · 被引用 4 次
- Extreme Classification via Adversarial Softmax ApproximationRobert Bamler, Stephan MandtICLR 2020 · 被引用 25 次
