A Cognitive Regularizer for Language Modeling
Jason Wei, Clara Meister, Ryan Cotterell
Abstract
The uniform information density (UID) hypothesis, which posits that speakers behaving optimally tend to distribute information uniformly across a linguistic signal, has gained traction in psycholinguistics as an explanation for certain syntactic, morphological, and prosodic choices. In this work, we explore whether the UID hypothesis can be operationalized as an inductive bias for statistical language modeling. Specifically, we augment the canonical MLE objective for training language models with a regularizer that encodes UID. In experiments on ten languages spanning five language families, we find that using UID regularization consistently improves perplexity in language models, having a larger effect when training data is limited. Moreover, via an analysis of generated sequences, we find that UID-regularized language models have other desirable properties, e.g., they generate text that is more lexically diverse. Our results not only suggest that UID is a reasonable inductive bias for language modeling, but also provide an alternative validation of the UID hypothesis using modern-day NLP tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8f9dcc1-4c71-4bea-b750-5139c05dc0deCited by top-tier papers3
- Information Value: Measuring Utterance Predictability as Distance from Plausible AlternativesMario Giulianelli, Sarenne Wallbridge, Raquel FernándezEMNLP 2023 · 4 citations
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger et al.EMNLP 2021 · 4 citations
- Lower Perplexity is Not Always Human-LikeTatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida et al.ACL 2021
Builds on3
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 26 citations
- Generalized Entropy Regularization or: There's Nothing Special about Label SmoothingClara Meister, Elizabeth Salesky, Ryan CotterellACL 2020 · 4 citations
Related papers
- Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form DiscourseEleftheria Tsipidi, Franz Nowak, Ryan Cotterell, Ethan Wilcox et al.EMNLP 2024 · 2 citations
- Uniform Information Density and Syntactic Reduction: Revisiting that-Mentioning in English Complement ClausesHailin Hao, Elsi KaiserEMNLP 2025
- The Harmonic Structure of Information ContoursEleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu et al.ACL 2025 · 6 citations
- Expect the Unexpected? Testing the Surprisal of Salient EntitiesJessica Lin, Amir ZeldesACL 2026 · 2 citations
- Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic SmoothingRichard Diehl Martinez, Zébulon Goriely, Andrew Caines, Paula Buttery et al.EMNLP 2024 · 1 citation
