Softmax Bottleneck Makes Language Models Unable to Represent Multi-mode Word Distributions
Haw-Shiuan Chang, Andrew McCallum
Abstract
Neural language models (LMs) such as GPT-2 estimate the probability distribution over the next word by a softmax over the vocabulary. The softmax layer produces the distribution based on the dot products of a single hidden state and the embeddings of words in the vocabulary. However, we discover that this single hidden state cannot produce all probability distributions regardless of the LM size or training data size because the single hidden state embedding cannot be close to the embeddings of all the possible next words simultaneously when there are other interfering word embeddings between them. In this work, we demonstrate the importance of this limitation both theoretically and practically. Our work not only deepens our understanding of softmax bottleneck and mixture of softmax (MoS) but also inspires us to propose multi-facet softmax (MFS) to address the limitations of MoS. Extensive empirical analyses confirm our findings and show that against MoS, the proposed MFS achieves two-fold improvements in the perplexity of GPT-2 and BERT. "The greater the ambiguity, the greater the pleasure." -Milan Kundera
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6c1bc39-918c-405a-bd7a-bb942a23db63Cited by top-tier papers10
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta et al.ICLR 2024 · 31 citations
- Sequences of Logits Reveal the Low Rank Structure of Language ModelsNoah Golowich, Allen Liu, Abhishek ShettyICLR 2026 · 9 citations
- Quantum Doubly Stochastic TransformersJannis Born, Filip Skogh, Kahn Rhrissorrakrai, Filippo Utro et al.NeurIPS 2025 · 7 citations
- HistAlign: Improving Context Dependency in Language Generation by Aligning with HistoryDavid Wan, Shiyue Zhang, Mohit BansalEMNLP 2023 · 4 citations
- Chatgpt Inaccuracy Mitigation During Technical Report Understanding: Are we There Yet?Salma Begum Tamanna, Gias Uddin, Song Wang, Lan Xia et al.ICSE 2025 · 2 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama et al.NeurIPS 2022 · 349 citations
- A Mutual Information Maximization Perspective of Language Representation LearningLingpeng Kong, Cyprien de Masson d'Autume, Lei Yu, Wang Ling et al.ICLR 2020 · 179 citations
Related papers
- On the Softmax Bottleneck of Recurrent Language ModelsDwarak Govind Parthiban, Yongyi Mao, Diana InkpenAAAI 2021 · 3 citations
- The Softmax Bottleneck Does Not Limit the Probabilities of the Most Likely TokensRonen Basri, David JacobsICLR 2026
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 2 citations
- Deriving Neural Scaling Laws from the Statistics of Natural LanguageFrancesco Cagnetta, Allan Raventos, Surya Ganguli, Matthieu WyartICML 2026
- Universal One-third Time Scaling in Learning Peaked DistributionsYizhou Liu, Ziming Liu, Cengiz Pehlevan, Jeff GoreICML 2026 · 6 citations
