Are Language Models Any Good at Density Modeling?
Sriram Ranga, Sai Shashank Bedampeta, Rui Mao, Anupam Chattopadhyay
摘要
Large Language Models (LLMs) surprised the world with their ability to mimic humans in writing and are starting to be used as simulations of human writers for various kinds of linguistic analyses. However, these analyses rest on the belief that LLMs are good density models that accurately capture the underlying probability distribution of the language. In this paper, we question this basic assumption and try to evaluate language models on their density modelling capabilities. Since a ground truth does not exist for the probability distribution of any natural language, we come up with a synthetic language made up of decimal numbers written in words in English. We train language models from scratch on various probability distributions over this synthetic language and compare the distributions learned by the models with the original distributions. Experiments show that language models can learn underlying probability distributions across a wide range of cases, but they fail when those distributions depend on deep semantic properties of numbers that cannot be inferred from syntactic patterns. Additionally, we observed a strong bias in the models towards numbers that frequently occur as substrings within other numbers. This suggests that such a bias possibly exists in real-world natural language models as well, and negatively impacts downstream tasks and analyses that rely on model-generated probabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Prompting is not a substitute for probability measurements in large language modelsJennifer Hu, Roger LevyEMNLP 2023 · 被引用 31 次
- Towards Understanding the Characteristics of Code Generation Errors Made by Large Language ModelsZhijie Wang, Zijie Zhou, Da Song, Yuheng Huang 等ICSE 2025 · 被引用 12 次
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger 等EMNLP 2021 · 被引用 4 次
- Language Model Evaluation Beyond PerplexityClara Meister, Ryan CotterellACL 2021
- Enough Coin Flips Can Make LLMs Act BayesianRitwik Gupta, Rodolfo Corona, Jiaxin Ge, Eric Wang 等ACL 2025
相关 Paper
- Language Models Learn Universal Representations of Numbers and Here's Why You Should CareMichal Stefánik, Timothee Mickus, Marek Kadlcík, Bertram Højer 等ACL 2026 · 被引用 1 次
- Benford's Curse: Tracing Digit Bias to Numerical Hallucination in LLMsJiandong Shao, Yao Lu, Jianfei YangNeurIPS 2025 · 被引用 8 次
- Eliciting Numerical Predictive Distributions of LLMs Without Auto-RegressionJulianna Piskorz, Kasia Kobalczyk, Mihaela van der SchaarICLR 2026 · 被引用 2 次
- What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular LanguagesNadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda 等ACL 2024
- Evaluating Distributional Distortion in Neural Language ModelingBenjamin LeBrun, Alessandro Sordoni, Timothy J. O'DonnellICLR 2022 · 被引用 26 次
