Norm of Word Embedding Encodes Information Gain
Momose Oyama, Sho Yokoi, Hidetoshi Shimodaira
Abstract
Distributed representations of words encode lexical semantic information, but what type of information is encoded and how? Focusing on the skip-gram with negative-sampling method, we found that the squared norm of static word embedding encodes the information gain conveyed by the word; the information gain is defined by the Kullback-Leibler divergence of the co-occurrence distribution of the word to the unigram distribution. Our findings are explained by the theoretical framework of the exponential family of probability distributions and confirmed through precise experiments that remove spurious correlations arising from word frequency. This theory also extends to contextualized word embeddings in language models or any neural networks with the softmax output layer. We also demonstrate that both the KL divergence and the squared norm of embedding provide a useful metric of the informativeness of a word in tasks such as keyword extraction, proper-noun discrimination, and hypernym discrimination.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42f4ea26-8a80-45df-8348-d548d19779acCited by top-tier papers11
- Latent Space Translation via Semantic AlignmentValentino Maiorca, Luca Moschella, Antonio Norelli, Marco Fumero et al.NeurIPS 2023 · 59 citations
- Coding-PTMs: How to Find Optimal Code Pre-trained Models for Code Embedding in Vulnerability Detection?Yu Zhao, Lina Gong, Zhiqiu Huang, Yongwei Wang et al.ASE 2024 · 10 citations
- Learning is Forgetting; LLM Training As Lossy CompressionHenry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen et al.ICLR 2026 · 6 citations
- Revisiting Anisotropy in Language Transformers: The Geometry of Learning DynamicsRaphael Bernas, Fanny Jourdan, Antonin Poché, Céline HudelotICML 2026 · 3 citations
- Zipfian WhiteningSho Yokoi, Han Bao, Hiroto Kurita, Hidetoshi ShimodairaNeurIPS 2024 · 3 citations
Builds on1
Related papers
- Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding WordsJosef Klafka, Allyson EttingerACL 2020 · 2 citations
- Speakers Fill Lexical Semantic Gaps with ContextTiago Pimentel, Rowan Hall Maudslay, Damián E. Blasi, Ryan CotterellEMNLP 2020 · 1 citation
- Static Word Embeddings for Sentence Semantic RepresentationTakashi Wada, Yuki Hirakawa, Ryotaro Shimizu, Takahiro Kawashima et al.EMNLP 2025 · 1 citation
- Obtaining Better Static Word Embeddings Using Contextual Embedding ModelsPrakhar Gupta, Martin JaggiACL 2021
- Unified Interpretation of Softmax Cross-Entropy and Negative Sampling: With Case Study for Knowledge Graph EmbeddingHidetaka Kamigaito, Katsuhiko HayashiACL 2021
