On the Softmax Bottleneck of Recurrent Language Models
Dwarak Govind Parthiban, Yongyi Mao, Diana Inkpen
Abstract
Recent research has pointed to a limitation of word-level neural language models with softmax outputs. This limitation, known as the "softmax bottleneck" refers to the inability of these models to produce high-rank log probability (log P ) matrices. Various solutions have been proposed to break this bottleneck, including Mixture of Softmaxes, SigSoftmax, and Linear Monotonic Softmax with Piecewise Linear Increasing Functions. They were reported to offer better performance in terms of perplexity on test data. A natural perception from these results is a strong positive correlation between the rank of the log P matrix and the model's performance. In this work, we show via an extensive empirical study that such a correlation is fairly weak and that the high-rank of the log P matrix is neither necessary nor sufficient for better test perplexity. Although our results are empirical, they are established in part via the construction of a rich family of models, which we call Generalized SigSoftmax. They are able to create diverse ranks for the log P matrices. We also present an investigation as to why the proposed solutions achieve better performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5acb863-03fd-4d0d-bee3-60d86ee7d7a9Cited by top-tier papers2
- The Softmax Bottleneck Does Not Limit the Probabilities of the Most Likely TokensRonen Basri, David JacobsICLR 2026
- Softmax Bottleneck Makes Language Models Unable to Represent Multi-mode Word DistributionsHaw-Shiuan Chang, Andrew McCallumACL 2022
Builds on1
Related papers
- On the Theoretical Limitations of Embedding-based Link PredictionSamy Badreddine, Emile van Krieken, Luciano SerafiniICML 2026
- Universal One-third Time Scaling in Learning Peaked DistributionsYizhou Liu, Ziming Liu, Cengiz Pehlevan, Jeff GoreICML 2026 · 6 citations
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta et al.ICLR 2024 · 31 citations
- What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular LanguagesNadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda et al.ACL 2024
- A Tale of Two Perplexities: Sensitivity of Neural Language Models to Lexical Retrieval Deficits in Dementia of the Alzheimer's TypeTrevor Cohen, Serguei PakhomovACL 2020 · 2 citations
