Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
Richard Diehl Martinez, Zébulon Goriely, Andrew Caines, Paula Buttery, Lisa Beinborn
摘要
Language models strongly rely on frequency information because they maximize the likelihood of tokens during pre-training. As a consequence, language models tend not to generalize well to tokens that are seldom seen during training. Moreover, maximum likelihood training has been discovered to give rise to anisotropy: representations of tokens in a model tend to cluster tightly in a highdimensional cone, rather than spreading out over their representational capacity. Our work introduces a method for quantifying the frequency bias of a language model by assessing sentence-level perplexity with respect to token-level frequency. We then present a method for reducing the frequency bias of a language model by inducing a syntactic prior over token representations during pre-training. Our Syntactic Smoothing method adjusts the maximum likelihood objective function to distribute the learning signal to syntactically similar tokens. This approach results in better performance on infrequent English tokens and a decrease in anisotropy. We empirically show that the degree of anisotropy in a model correlates with its frequency bias. rdiehlmartinez/syntactic-smoothing
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Empirical Analysis of Decoding Biases in Masked Diffusion ModelsPengcheng Huang, Tianming Liu, Zhenghao Liu, Yukun Yan 等ACL 2026 · 被引用 15 次
- Enhancing Text-to-Image Diffusion Transformer via Split-Text ConditioningYu Zhang, Jialei Zhou, Xinchen Li, Qi Zhang 等NeurIPS 2025 · 被引用 11 次
- Developmentally-plausible Working Memory Shapes a Critical Period for Language AcquisitionMasato Mita, Ryo Yoshida, Yohei OsekiACL 2025 · 被引用 7 次
- Revisiting Anisotropy in Language Transformers: The Geometry of Learning DynamicsRaphael Bernas, Fanny Jourdan, Antonin Poché, Céline HudelotICML 2026 · 被引用 3 次
- Disentangling Geometry, Performance, and Training in Language ModelsAtharva Kulkarni, Jacob Mitchell Springer, Arjun Subramonian, Swabha SwayamdiptaICML 2026 · 被引用 1 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
- Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive MimickingTimo Schick, Hinrich SchützeAAAI 2020 · 被引用 106 次
相关 Paper
- Frequency Effects on Syntactic Rule Learning in TransformersJason Wei, Dan Garrette, Tal Linzen, Ellie PavlickEMNLP 2021 · 被引用 40 次
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-trainingWoojin Chung, Jeonghoon KimNeurIPS 2025 · 被引用 6 次
- Token-level Adaptive Training for Neural Machine TranslationShuhao Gu, Jinchao Zhang, Fandong Meng, Yang Feng 等EMNLP 2020 · 被引用 32 次
- Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language ModelsRobert Wolfe, Aylin CaliskanEMNLP 2021 · 被引用 27 次
- Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For PerplexityYeongbin Seo, Gayoung Kim, Jaehyung Kim, Jinyoung YeoICLR 2026
