Softmax Bottleneck Makes Language Models Unable to Represent Multi-mode Word Distributions
Haw-Shiuan Chang, Andrew McCallum
摘要
Neural language models (LMs) such as GPT-2 estimate the probability distribution over the next word by a softmax over the vocabulary. The softmax layer produces the distribution based on the dot products of a single hidden state and the embeddings of words in the vocabulary. However, we discover that this single hidden state cannot produce all probability distributions regardless of the LM size or training data size because the single hidden state embedding cannot be close to the embeddings of all the possible next words simultaneously when there are other interfering word embeddings between them. In this work, we demonstrate the importance of this limitation both theoretically and practically. Our work not only deepens our understanding of softmax bottleneck and mixture of softmax (MoS) but also inspires us to propose multi-facet softmax (MFS) to address the limitations of MoS. Extensive empirical analyses confirm our findings and show that against MoS, the proposed MFS achieves two-fold improvements in the perplexity of GPT-2 and BERT. "The greater the ambiguity, the greater the pleasure." -Milan Kundera
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta 等ICLR 2024 · 被引用 31 次
- Sequences of Logits Reveal the Low Rank Structure of Language ModelsNoah Golowich, Allen Liu, Abhishek ShettyICLR 2026 · 被引用 9 次
- Quantum Doubly Stochastic TransformersJannis Born, Filip Skogh, Kahn Rhrissorrakrai, Filippo Utro 等NeurIPS 2025 · 被引用 7 次
- HistAlign: Improving Context Dependency in Language Generation by Aligning with HistoryDavid Wan, Shiyue Zhang, Mohit BansalEMNLP 2023 · 被引用 4 次
- Chatgpt Inaccuracy Mitigation During Technical Report Understanding: Are we There Yet?Salma Begum Tamanna, Gias Uddin, Song Wang, Lan Xia 等ICSE 2025 · 被引用 2 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama 等NeurIPS 2022 · 被引用 349 次
- A Mutual Information Maximization Perspective of Language Representation LearningLingpeng Kong, Cyprien de Masson d'Autume, Lei Yu, Wang Ling 等ICLR 2020 · 被引用 179 次
相关 Paper
- On the Softmax Bottleneck of Recurrent Language ModelsDwarak Govind Parthiban, Yongyi Mao, Diana InkpenAAAI 2021 · 被引用 3 次
- The Softmax Bottleneck Does Not Limit the Probabilities of the Most Likely TokensRonen Basri, David JacobsICLR 2026
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 被引用 2 次
- Deriving Neural Scaling Laws from the Statistics of Natural LanguageFrancesco Cagnetta, Allan Raventos, Surya Ganguli, Matthieu WyartICML 2026
- Universal One-third Time Scaling in Learning Peaked DistributionsYizhou Liu, Ziming Liu, Cengiz Pehlevan, Jeff GoreICML 2026 · 被引用 6 次
