Zipfian Whitening
Sho Yokoi, Han Bao, Hiroto Kurita, Hidetoshi Shimodaira
摘要
The word embedding space in neural models is skewed, and correcting this can improve task performance. We point out that most approaches for modeling, correcting, and measuring the symmetry of an embedding space implicitly assume that the word frequencies are uniform; in reality, word frequencies follow a highly non-uniform distribution, known as Zipf's law. Surprisingly, simply performing PCA whitening weighted by the empirical word frequency that follows Zipf's law significantly improves task performance, surpassing established baselines. From a theoretical perspective, both our approach and existing methods can be clearly categorized: word representations are distributed according to an exponential family with either uniform or Zipfian base measures. By adopting the latter approach, we can naturally emphasize informative low-frequency words in terms of their vector norm, which becomes evident from the information-geometric perspective, and in terms of the loss functions for imbalanced classification. Additionally, our theory corroborates that popular natural language processing methods, such as skip-gram negative sampling, WhiteningBERT, and headless language models, work well just because their word embeddings encode the empirical word frequency into the underlying probabilistic model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Distribution-Aware Exploration for Adaptive HNSW SearchChao Zhang, Renée J. MillerSIGMOD 2026 · 被引用 9 次
- SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale CorporaMasataka Yoneda, Yusuke Matsushita, Go Kamoda, Kohei Suenaga 等ICML 2026
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain 等ICLR 2021 · 被引用 937 次
- Graph Neural Networks Exponentially Lose Expressive Power for Node ClassificationKenta Oono, Taiji SuzukiICLR 2020 · 被引用 864 次
相关 Paper
- Cross-Domain Empirical Risk Minimization for Unbiased Long-Tailed ClassificationBeier Zhu, Yulei Niu, Xian-Sheng Hua, Hanwang ZhangAAAI 2022 · 被引用 50 次
- Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's LawFrederik Kunstner, Francis BachNeurIPS 2025 · 被引用 21 次
- Norm of Word Embedding Encodes Information GainMomose Oyama, Sho Yokoi, Hidetoshi ShimodairaEMNLP 2023 · 被引用 6 次
- A New Formulation of Zipf's Meaning-Frequency Law through Contextual DiversityRyo Nagata, Kumiko Tanaka-IshiiACL 2025
- The Power of Power Law: Asymmetry Enables Compositional ReasoningZixuan Wang, Xingyu Dang, Jason Lee, Kaifeng LyuICML 2026 · 被引用 1 次
