What Are the Odds? Language Models Are Capable of Probabilistic Reasoning
Akshay Paruchuri, Jake Garrison, Shun Liao, John Hernandez, Jacob E. Sunshine, Tim Althoff, Xin Liu, Daniel McDuff
Abstract
Language models (LM) are capable of remarkably complex linguistic tasks; however, numerical reasoning is an area in which they frequently struggle. An important but rarely evaluated form of reasoning is understanding probability distributions. In this paper, we focus on evaluating the probabilistic reasoning capabilities of LMs using idealized and real-world statistical distributions. We perform a systematic evaluation of state-of-the-art LMs on three tasks: estimating percentiles, drawing samples, and calculating probabilities. We evaluate three ways to provide context to LMs 1) anchoring examples from within a distribution or family of distributions, 2) real-world context, 3) summary statistics on which to base a Normal approximation. Models can make inferences about distributions, and can be further aided by the incorporation of real-world context, example shots and simplified assumptions, even if these assumptions are incorrect or misspecified. To conduct this work, we developed a comprehensive benchmark distribution dataset with associated question-answer pairs that we have released publicly. * Work completed during an internship at Google. ## Consider the following distribution: Type: Log-Normal Distribution Characteristics: This distribution models values that are the result of the multiplicative product of many independent random variables, such as income levels, stock prices, or city sizes. Log Mean (mu): 3.543 Log Sigma (sigma): 0.677 These parameters mean that the natural logarithm of the values follows a normal distribution with the speci�ed mean and standard deviation. Your task is to estimate the percentile of a given average exercise minutes count value for a population that regularly uses Fitbit devices and is active on a daily basis. The data is �ltered for individuals aged 18-65. The data is age-balanced and gender-balanced, and pe�ains to the U.S. population only. Consider the following parameters that describe a normal distribution:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Probabilistic Reasoning with LLMs for Privacy Risk EstimationJonathan Zheng, Alan Ritter, Sauvik Das, Wei (Coco) XuNeurIPS 2025 · 3 citations
- OpenEstimate: Evaluating LLMs on Reasoning Under Uncertainty with Real-World DataAlana Renda, Jillian Ross, Jacob AndreasICLR 2026 · 3 citations
- Bonsai: Interpretable Tree-Adaptive Grounded ReasoningKate Sanders, Benjamin Van DurmeAAAI 2026 · 1 citation
Builds on9
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 163 citations
- A Survey of Deep Learning for Mathematical ReasoningPan Lu, Liang Qiu, Wenhao Yu, Sean Welleck et al.ACL 2023 · 43 citations
Related papers
- Reasoning over Uncertain Text by Generative Large Language ModelsAliakbar Nafar, Kristen Brent Venable, Parisa KordjamshidiAAAI 2025 · 13 citations
- Are Language Models Any Good at Density Modeling?Sriram Ranga, Sai Shashank Bedampeta, Rui Mao, Anupam ChattopadhyayAAAI 2026
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 38 citations
- MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation DatasetWeiqi Wang, Yangqiu SongACL 2025
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language ModelsMihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura et al.ACL 2024
