The Softmax Bottleneck Does Not Limit the Probabilities of the Most Likely Tokens
Ronen Basri, David Jacobs
Abstract
In many popular transformer architectures, an output projection matrix linearly maps lower-dimensional embeddings into a higher-dimensional space of logits. It has been shown that this leads to a softmax bottleneck that prevents the production of arbitrary probability distributions. It has been argued that this limits large language models (LLMs) in their ability to express next token probabilities that perfectly align with the statistics of natural language. We focus on the ability of such models to produce accurate probabilities for just the top- tokens. We provide theoretical bounds that show that even a randomly initialized projection matrix can successfully do this for rather large values of , supported by empirical results on both random and trained matrices. This raises questions about whether the softmax bottleneck significantly limits the capabilities of LLMs. We also derive bounds on the maximum number of probabilities that any trained output projection matrix can specify.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00226bbe-7eb0-41e0-aa99-d059b254d4c5Builds on9
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
- The Expressive Power of Transformers with Chain of ThoughtWilliam Merrill, Ashish SabharwalICLR 2024 · 243 citations
- Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?Tokio Kajitsuka, Issei SatoICLR 2024 · 31 citations
Related papers
- On the Softmax Bottleneck of Recurrent Language ModelsDwarak Govind Parthiban, Yongyi Mao, Diana InkpenAAAI 2021 · 3 citations
- Softmax Bottleneck Makes Language Models Unable to Represent Multi-mode Word DistributionsHaw-Shiuan Chang, Andrew McCallumACL 2022
- Universal One-third Time Scaling in Learning Peaked DistributionsYizhou Liu, Ziming Liu, Cengiz Pehlevan, Jeff GoreICML 2026 · 6 citations
- Low-Rank Softmax Can Have Unargmaxable Classes in Theory but Rarely in PracticeAndreas Grivas, Nikolay Bogoychev, Adam LopezACL 2022
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta et al.ICLR 2024 · 31 citations
