Norm of Word Embedding Encodes Information Gain
Momose Oyama, Sho Yokoi, Hidetoshi Shimodaira
摘要
Distributed representations of words encode lexical semantic information, but what type of information is encoded and how? Focusing on the skip-gram with negative-sampling method, we found that the squared norm of static word embedding encodes the information gain conveyed by the word; the information gain is defined by the Kullback-Leibler divergence of the co-occurrence distribution of the word to the unigram distribution. Our findings are explained by the theoretical framework of the exponential family of probability distributions and confirmed through precise experiments that remove spurious correlations arising from word frequency. This theory also extends to contextualized word embeddings in language models or any neural networks with the softmax output layer. We also demonstrate that both the KL divergence and the squared norm of embedding provide a useful metric of the informativeness of a word in tasks such as keyword extraction, proper-noun discrimination, and hypernym discrimination.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Latent Space Translation via Semantic AlignmentValentino Maiorca, Luca Moschella, Antonio Norelli, Marco Fumero 等NeurIPS 2023 · 被引用 59 次
- Coding-PTMs: How to Find Optimal Code Pre-trained Models for Code Embedding in Vulnerability Detection?Yu Zhao, Lina Gong, Zhiqiu Huang, Yongwei Wang 等ASE 2024 · 被引用 10 次
- Learning is Forgetting; LLM Training As Lossy CompressionHenry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen 等ICLR 2026 · 被引用 6 次
- Revisiting Anisotropy in Language Transformers: The Geometry of Learning DynamicsRaphael Bernas, Fanny Jourdan, Antonin Poché, Céline HudelotICML 2026 · 被引用 3 次
- Zipfian WhiteningSho Yokoi, Han Bao, Hiroto Kurita, Hidetoshi ShimodairaNeurIPS 2024 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding WordsJosef Klafka, Allyson EttingerACL 2020 · 被引用 2 次
- Speakers Fill Lexical Semantic Gaps with ContextTiago Pimentel, Rowan Hall Maudslay, Damián E. Blasi, Ryan CotterellEMNLP 2020 · 被引用 1 次
- Static Word Embeddings for Sentence Semantic RepresentationTakashi Wada, Yuki Hirakawa, Ryotaro Shimizu, Takahiro Kawashima 等EMNLP 2025 · 被引用 1 次
- Obtaining Better Static Word Embeddings Using Contextual Embedding ModelsPrakhar Gupta, Martin JaggiACL 2021
- Unified Interpretation of Softmax Cross-Entropy and Negative Sampling: With Case Study for Knowledge Graph EmbeddingHidetaka Kamigaito, Katsuhiko HayashiACL 2021
