Rankmax: An Adaptive Projection Alternative to the Softmax Function
Weiwei Kong, Walid Krichene, Nicolas Mayoraz, Steffen Rendle, Li Zhang
Abstract
Many machine learning models involve mapping a score vector to a probability vector. Usually, this is done by projecting the score vector onto a probability simplex, and such projections are often characterized as Lipschitz continuous approximations of the argmax function, whose Lipschitz constant is controlled by a parameter that is similar to a softmax temperature. The aforementioned parameter has been observed to affect the quality of these models and is typically either treated as a constant or decayed over time. In this work, we propose a method that adapts this parameter to individual training examples. The resulting method exhibits desirable properties, such as sparsity of its support and numerically efficient implementation, and we find that it significantly outperforms competing non-adaptive projection methods. In our analysis, we also derive the general solution of (Bregman) projections onto the (n, k)-simplex, a result which may be of independent interest.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63fac855-ef54-4357-bee4-dc57ab349fefCited by top-tier papers4
- Learning Recommender Systems with Implicit Feedback via Soft Target EnhancementMingyue Cheng, Fajie Yuan, Qi Liu, Shenyang Ge et al.SIGIR 2021 · 22 citations
- Supervised Tree-Wasserstein DistanceYuki Takezawa, Ryoma Sato, Makoto YamadaICML 2021 · 14 citations
- Listwise Learning to Rank Based on Approximate Rank IndicatorsThibaut Thonet, Yagmur Gizem Cinar, Éric Gaussier, Minghan Li et al.AAAI 2022 · 12 citations
- Sparse Communication via Mixed DistributionsAntónio Farinhas, Wilker Aziz, Vlad Niculae, André F. T. MartinsICLR 2022 · 3 citations
Related papers
- MultiMax: Sparse and Multi-Modal Attention LearningYuxuan Zhou, Mario Fritz, Margret KeuperICML 2024 · 4 citations
- Softmax is not Enough (for Sharp Size Generalisation)Petar Velickovic, Christos Perivolaropoulos, Federico Barbero, Razvan PascanuICML 2025
- Adaptive Sampling for Efficient Softmax ApproximationTavor Z. Baharav, Ryan Kang, Colin Sullivan, Mo Tiwari et al.NeurIPS 2024 · 7 citations
- Binary Hypothesis Testing for Softmax Models and Leverage Score ModelsYuzhou Gu, Zhao Song, Junze YinICML 2025
- Escaping the Gravitational Pull of SoftmaxJincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li et al.NeurIPS 2020 · 56 citations
