Evaluating Distributional Distortion in Neural Language Modeling
Benjamin LeBrun, Alessandro Sordoni, Timothy J. O'Donnell
Abstract
A fundamental characteristic of natural language is the high rate at which speakers produce novel expressions. Because of this novelty, a heavy-tail of rare events accounts for a significant amount of the total probability mass of distributions in language (Baayen, 2001). Standard language modeling metrics such as perplexity quantify the performance of language models (LM) in aggregate. As a result, we have relatively little understanding of whether neural LMs accurately estimate the probability of sequences in this heavy-tail of rare events. To address this gap, we develop a controlled evaluation scheme which uses generative models trained on natural data as artificial languages from which we can exactly compute sequence probabilities. Training LMs on generations from these artificial languages, we compare the sequence-level probability estimates given by LMs to the true probabilities in the target language. Our experiments reveal that LSTM and Transformer language models (i) systematically underestimate the probability of sequences drawn from the target language, and (ii) do so more severely for less-probable sequences. Investigating where this probability mass went, (iii) we find that LMs tend to overestimate the probability of ill formed (perturbed) sequences. In addition, we find that this underestimation behaviour (iv) is weakened, but not eliminated by greater amounts of training data, and (v) is exacerbated for target distributions with lower entropy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2a5ce85-48c8-4ebd-9ecc-cee6714e3f3eCited by top-tier papers13
- Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context LearningXinyi Wang, Wanrong Zhu, Michael Saxon, Mark Steyvers et al.NeurIPS 2023 · 206 citations
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton et al.ICML 2024 · 123 citations
- The Best Instruction-Tuning Data are Those That FitDylan Zhang, Qirun Dai, Hao PengNeurIPS 2025 · 59 citations
- Emergent Representations of Program Semantics in Language Models Trained on ProgramsCharles Jin, Martin C. RinardICML 2024 · 34 citations
- Why Less is More (Sometimes): A Theory of Data CurationElvis Dohmatob, Mohammad Pezeshki, Reyhane Askari HemmatICLR 2026 · 11 citations
Builds on7
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 436 citations
- Consistency of a Recurrent Language Model With Respect to Incomplete DecodingSean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang et al.EMNLP 2020 · 37 citations
- Moving Down the Long Tail of Word Sense Disambiguation with Gloss Informed Bi-encodersTerra Blevins, Luke ZettlemoyerACL 2020 · 19 citations
- Mogrifier LSTMGábor Melis, Tomás Kociský, Phil BlunsomICLR 2020
Related papers
- Rare Event Analysis of Large Language ModelsJake McAllister Dorman, Edward Gillman, Dominic C Rose, Jamie Mair et al.ICML 2026
- Calibration, Entropy Rates, and Memory in Language ModelsMark Braverman, Xinyi Chen, Sham M. Kakade, Karthik Narasimhan et al.ICML 2020 · 48 citations
- Are Language Models Any Good at Density Modeling?Sriram Ranga, Sai Shashank Bedampeta, Rui Mao, Anupam ChattopadhyayAAAI 2026
- Language Model Evaluation Beyond PerplexityClara Meister, Ryan CotterellACL 2021
- Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can ProduceHaojin Wang, Zining Zhu, Freda ShiEMNLP 2025
