Zipfian Whitening
Sho Yokoi, Han Bao, Hiroto Kurita, Hidetoshi Shimodaira
Abstract
The word embedding space in neural models is skewed, and correcting this can improve task performance. We point out that most approaches for modeling, correcting, and measuring the symmetry of an embedding space implicitly assume that the word frequencies are uniform; in reality, word frequencies follow a highly non-uniform distribution, known as Zipf's law. Surprisingly, simply performing PCA whitening weighted by the empirical word frequency that follows Zipf's law significantly improves task performance, surpassing established baselines. From a theoretical perspective, both our approach and existing methods can be clearly categorized: word representations are distributed according to an exponential family with either uniform or Zipfian base measures. By adopting the latter approach, we can naturally emphasize informative low-frequency words in terms of their vector norm, which becomes evident from the information-geometric perspective, and in terms of the loss functions for imbalanced classification. Additionally, our theory corroborates that popular natural language processing methods, such as skip-gram negative sampling, WhiteningBERT, and headless language models, work well just because their word embeddings encode the empirical word frequency into the underlying probabilistic model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Distribution-Aware Exploration for Adaptive HNSW SearchChao Zhang, Renée J. MillerSIGMOD 2026 · 9 citations
- SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale CorporaMasataka Yoneda, Yusuke Matsushita, Go Kamoda, Kohei Suenaga et al.ICML 2026
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- Graph Neural Networks Exponentially Lose Expressive Power for Node ClassificationKenta Oono, Taiji SuzukiICLR 2020 · 864 citations
Related papers
- Cross-Domain Empirical Risk Minimization for Unbiased Long-Tailed ClassificationBeier Zhu, Yulei Niu, Xian-Sheng Hua, Hanwang ZhangAAAI 2022 · 50 citations
- Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's LawFrederik Kunstner, Francis BachNeurIPS 2025 · 21 citations
- Norm of Word Embedding Encodes Information GainMomose Oyama, Sho Yokoi, Hidetoshi ShimodairaEMNLP 2023 · 6 citations
- A New Formulation of Zipf's Meaning-Frequency Law through Contextual DiversityRyo Nagata, Kumiko Tanaka-IshiiACL 2025
- The Power of Power Law: Asymmetry Enables Compositional ReasoningZixuan Wang, Xingyu Dang, Jason Lee, Kaifeng LyuICML 2026 · 1 citation
