Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
Richard Diehl Martinez, Zébulon Goriely, Andrew Caines, Paula Buttery, Lisa Beinborn
Abstract
Language models strongly rely on frequency information because they maximize the likelihood of tokens during pre-training. As a consequence, language models tend not to generalize well to tokens that are seldom seen during training. Moreover, maximum likelihood training has been discovered to give rise to anisotropy: representations of tokens in a model tend to cluster tightly in a highdimensional cone, rather than spreading out over their representational capacity. Our work introduces a method for quantifying the frequency bias of a language model by assessing sentence-level perplexity with respect to token-level frequency. We then present a method for reducing the frequency bias of a language model by inducing a syntactic prior over token representations during pre-training. Our Syntactic Smoothing method adjusts the maximum likelihood objective function to distribute the learning signal to syntactically similar tokens. This approach results in better performance on infrequent English tokens and a decrease in anisotropy. We empirically show that the degree of anisotropy in a model correlates with its frequency bias. rdiehlmartinez/syntactic-smoothing
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d629b7f4-9e1d-4730-86e3-75f1abddb2edCited by top-tier papers5
- Empirical Analysis of Decoding Biases in Masked Diffusion ModelsPengcheng Huang, Tianming Liu, Zhenghao Liu, Yukun Yan et al.ACL 2026 · 15 citations
- Enhancing Text-to-Image Diffusion Transformer via Split-Text ConditioningYu Zhang, Jialei Zhou, Xinchen Li, Qi Zhang et al.NeurIPS 2025 · 11 citations
- Developmentally-plausible Working Memory Shapes a Critical Period for Language AcquisitionMasato Mita, Ryo Yoshida, Yohei OsekiACL 2025 · 7 citations
- Revisiting Anisotropy in Language Transformers: The Geometry of Learning DynamicsRaphael Bernas, Fanny Jourdan, Antonin Poché, Céline HudelotICML 2026 · 3 citations
- Disentangling Geometry, Performance, and Training in Language ModelsAtharva Kulkarni, Jacob Mitchell Springer, Arjun Subramonian, Swabha SwayamdiptaICML 2026 · 1 citation
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive MimickingTimo Schick, Hinrich SchützeAAAI 2020 · 106 citations
Related papers
- Frequency Effects on Syntactic Rule Learning in TransformersJason Wei, Dan Garrette, Tal Linzen, Ellie PavlickEMNLP 2021 · 40 citations
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-trainingWoojin Chung, Jeonghoon KimNeurIPS 2025 · 6 citations
- Token-level Adaptive Training for Neural Machine TranslationShuhao Gu, Jinchao Zhang, Fandong Meng, Yang Feng et al.EMNLP 2020 · 32 citations
- Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language ModelsRobert Wolfe, Aylin CaliskanEMNLP 2021 · 27 citations
- Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For PerplexityYeongbin Seo, Gayoung Kim, Jaehyung Kim, Jinyoung YeoICLR 2026
