A Cognitive Regularizer for Language Modeling
Jason Wei, Clara Meister, Ryan Cotterell
摘要
The uniform information density (UID) hypothesis, which posits that speakers behaving optimally tend to distribute information uniformly across a linguistic signal, has gained traction in psycholinguistics as an explanation for certain syntactic, morphological, and prosodic choices. In this work, we explore whether the UID hypothesis can be operationalized as an inductive bias for statistical language modeling. Specifically, we augment the canonical MLE objective for training language models with a regularizer that encodes UID. In experiments on ten languages spanning five language families, we find that using UID regularization consistently improves perplexity in language models, having a larger effect when training data is limited. Moreover, via an analysis of generated sequences, we find that UID-regularized language models have other desirable properties, e.g., they generate text that is more lexically diverse. Our results not only suggest that UID is a reasonable inductive bias for language modeling, but also provide an alternative validation of the UID hypothesis using modern-day NLP tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Information Value: Measuring Utterance Predictability as Distance from Plausible AlternativesMario Giulianelli, Sarenne Wallbridge, Raquel FernándezEMNLP 2023 · 被引用 4 次
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger 等EMNLP 2021 · 被引用 4 次
- Lower Perplexity is Not Always Human-LikeTatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida 等ACL 2021
它引用的顶会 Paper3
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 被引用 26 次
- Generalized Entropy Regularization or: There's Nothing Special about Label SmoothingClara Meister, Elizabeth Salesky, Ryan CotterellACL 2020 · 被引用 4 次
相关 Paper
- Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form DiscourseEleftheria Tsipidi, Franz Nowak, Ryan Cotterell, Ethan Wilcox 等EMNLP 2024 · 被引用 2 次
- Uniform Information Density and Syntactic Reduction: Revisiting that-Mentioning in English Complement ClausesHailin Hao, Elsi KaiserEMNLP 2025
- The Harmonic Structure of Information ContoursEleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu 等ACL 2025 · 被引用 6 次
- Expect the Unexpected? Testing the Surprisal of Salient EntitiesJessica Lin, Amir ZeldesACL 2026 · 被引用 2 次
- Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic SmoothingRichard Diehl Martinez, Zébulon Goriely, Andrew Caines, Paula Buttery 等EMNLP 2024 · 被引用 1 次
