Lune

ICLR2026Top-tier venue

The Softmax Bottleneck Does Not Limit the Probabilities of the Most Likely Tokens

Ronen Basri, David Jacobs

2026Year

Abstract

In many popular transformer architectures, an output projection matrix linearly maps lower-dimensional embeddings into a higher-dimensional space of logits. It has been shown that this leads to a softmax bottleneck that prevents the production of arbitrary probability distributions. It has been argued that this limits large language models (LLMs) in their ability to express next token probabilities that perfectly align with the statistics of natural language. We focus on the ability of such models to produce accurate probabilities for just the top-mm tokens. We provide theoretical bounds that show that even a randomly initialized projection matrix can successfully do this for rather large values of mm, supported by empirical results on both random and trained matrices. This raises questions about whether the softmax bottleneck significantly limits the capabilities of LLMs. We also derive bounds on the maximum number of probabilities that any trained output projection matrix can specify.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 00226bbe-7eb0-41e0-aa99-d059b254d4c5

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines