Softmax is not Enough (for Sharp Size Generalisation)
Petar Velickovic, Christos Perivolaropoulos, Federico Barbero, Razvan Pascanu
Abstract
A key property of reasoning systems is the ability to make sharp decisions on their input data. For contemporary AI systems, a key carrier of sharp behaviour is the softmax function, with its capability to perform differentiable query-key lookups. It is a common belief that the predictive power of networks leveraging softmax arises from "circuits" which sharply perform certain kinds of computations consistently across many diverse inputs. However, for these circuits to be robust, they would need to generalise well to arbitrary valid inputs. In this paper, we dispel this myth: even for tasks as simple as finding the maximum key, any learned circuitry must disperse as the number of items grows at test time. We attribute this to a fundamental limitation of the softmax function to robustly approximate sharp functions with increasing problem size, prove this phenomenon theoretically, and propose adaptive temperature as an ad-hoc technique for improving the sharpness of softmax at inference time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation ModelJingang QU, David Holzmüller, Gael Varoquaux, Marine Le MorvanICML 2026 · 85 citations
- Long-Context Generalization with Sparse AttentionPavlo Vasylenko, Hugo Pitorro, Andre F. T. Martins, Marcos V. TrevisoICLR 2026 · 19 citations
- What Happens Next? Anticipating Future Motion by Generating Point TrajectoriesGabrijel Boduljak, Laurynas Karazija, Iro Laina, Christian Rupprecht et al.ICLR 2026 · 10 citations
- Attention Smoothing Is All You Need For UnlearningSaleh Zare Zade, Xiangyu Zhou, Sijia Liu, Dongxiao ZhuICLR 2026 · 7 citations
- AdaSplash-2: Faster Differentiable Sparse AttentionNuno M. T. Gonçalves, Hugo Pitorro, Vlad Niculae, Edoardo Ponti et al.ICML 2026 · 3 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 1,717 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
Related papers
- Adaptive Sampling for Efficient Softmax ApproximationTavor Z. Baharav, Ryan Kang, Colin Sullivan, Mo Tiwari et al.NeurIPS 2024 · 7 citations
- Rankmax: An Adaptive Projection Alternative to the Softmax FunctionWeiwei Kong, Walid Krichene, Nicolas Mayoraz, Steffen Rendle et al.NeurIPS 2020 · 23 citations
- Binary Hypothesis Testing for Softmax Models and Leverage Score ModelsYuzhou Gu, Zhao Song, Junze YinICML 2025
- Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High DimensionsSamet Demir, Zafer DoganICML 2026 · 1 citation
- Enhancing Classifier Conservativeness and Robustness by PolynomialityZiqi Wang, Marco LoogCVPR 2022 · 2 citations
