Rankmax: An Adaptive Projection Alternative to the Softmax Function
Weiwei Kong, Walid Krichene, Nicolas Mayoraz, Steffen Rendle, Li Zhang
摘要
Many machine learning models involve mapping a score vector to a probability vector. Usually, this is done by projecting the score vector onto a probability simplex, and such projections are often characterized as Lipschitz continuous approximations of the argmax function, whose Lipschitz constant is controlled by a parameter that is similar to a softmax temperature. The aforementioned parameter has been observed to affect the quality of these models and is typically either treated as a constant or decayed over time. In this work, we propose a method that adapts this parameter to individual training examples. The resulting method exhibits desirable properties, such as sparsity of its support and numerically efficient implementation, and we find that it significantly outperforms competing non-adaptive projection methods. In our analysis, we also derive the general solution of (Bregman) projections onto the (n, k)-simplex, a result which may be of independent interest.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning Recommender Systems with Implicit Feedback via Soft Target EnhancementMingyue Cheng, Fajie Yuan, Qi Liu, Shenyang Ge 等SIGIR 2021 · 被引用 22 次
- Supervised Tree-Wasserstein DistanceYuki Takezawa, Ryoma Sato, Makoto YamadaICML 2021 · 被引用 14 次
- Listwise Learning to Rank Based on Approximate Rank IndicatorsThibaut Thonet, Yagmur Gizem Cinar, Éric Gaussier, Minghan Li 等AAAI 2022 · 被引用 12 次
- Sparse Communication via Mixed DistributionsAntónio Farinhas, Wilker Aziz, Vlad Niculae, André F. T. MartinsICLR 2022 · 被引用 3 次
相关 Paper
- MultiMax: Sparse and Multi-Modal Attention LearningYuxuan Zhou, Mario Fritz, Margret KeuperICML 2024 · 被引用 4 次
- Softmax is not Enough (for Sharp Size Generalisation)Petar Velickovic, Christos Perivolaropoulos, Federico Barbero, Razvan PascanuICML 2025
- Adaptive Sampling for Efficient Softmax ApproximationTavor Z. Baharav, Ryan Kang, Colin Sullivan, Mo Tiwari 等NeurIPS 2024 · 被引用 7 次
- Binary Hypothesis Testing for Softmax Models and Leverage Score ModelsYuzhou Gu, Zhao Song, Junze YinICML 2025
- Escaping the Gravitational Pull of SoftmaxJincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li 等NeurIPS 2020 · 被引用 56 次
