Evaluating Distributional Distortion in Neural Language Modeling
Benjamin LeBrun, Alessandro Sordoni, Timothy J. O'Donnell
摘要
A fundamental characteristic of natural language is the high rate at which speakers produce novel expressions. Because of this novelty, a heavy-tail of rare events accounts for a significant amount of the total probability mass of distributions in language (Baayen, 2001). Standard language modeling metrics such as perplexity quantify the performance of language models (LM) in aggregate. As a result, we have relatively little understanding of whether neural LMs accurately estimate the probability of sequences in this heavy-tail of rare events. To address this gap, we develop a controlled evaluation scheme which uses generative models trained on natural data as artificial languages from which we can exactly compute sequence probabilities. Training LMs on generations from these artificial languages, we compare the sequence-level probability estimates given by LMs to the true probabilities in the target language. Our experiments reveal that LSTM and Transformer language models (i) systematically underestimate the probability of sequences drawn from the target language, and (ii) do so more severely for less-probable sequences. Investigating where this probability mass went, (iii) we find that LMs tend to overestimate the probability of ill formed (perturbed) sequences. In addition, we find that this underestimation behaviour (iv) is weakened, but not eliminated by greater amounts of training data, and (v) is exacerbated for target distributions with lower entropy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context LearningXinyi Wang, Wanrong Zhu, Michael Saxon, Mark Steyvers 等NeurIPS 2023 · 被引用 206 次
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton 等ICML 2024 · 被引用 123 次
- The Best Instruction-Tuning Data are Those That FitDylan Zhang, Qirun Dai, Hao PengNeurIPS 2025 · 被引用 59 次
- Emergent Representations of Program Semantics in Language Models Trained on ProgramsCharles Jin, Martin C. RinardICML 2024 · 被引用 34 次
- Why Less is More (Sometimes): A Theory of Data CurationElvis Dohmatob, Mohammad Pezeshki, Reyhane Askari HemmatICLR 2026 · 被引用 11 次
它引用的顶会 Paper7
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 被引用 436 次
- Consistency of a Recurrent Language Model With Respect to Incomplete DecodingSean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang 等EMNLP 2020 · 被引用 37 次
- Moving Down the Long Tail of Word Sense Disambiguation with Gloss Informed Bi-encodersTerra Blevins, Luke ZettlemoyerACL 2020 · 被引用 19 次
- Mogrifier LSTMGábor Melis, Tomás Kociský, Phil BlunsomICLR 2020
相关 Paper
- Rare Event Analysis of Large Language ModelsJake McAllister Dorman, Edward Gillman, Dominic C Rose, Jamie Mair 等ICML 2026
- Calibration, Entropy Rates, and Memory in Language ModelsMark Braverman, Xinyi Chen, Sham M. Kakade, Karthik Narasimhan 等ICML 2020 · 被引用 48 次
- Are Language Models Any Good at Density Modeling?Sriram Ranga, Sai Shashank Bedampeta, Rui Mao, Anupam ChattopadhyayAAAI 2026
- Language Model Evaluation Beyond PerplexityClara Meister, Ryan CotterellACL 2021
- Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can ProduceHaojin Wang, Zining Zhu, Freda ShiEMNLP 2025
