Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
Minda Zhao, Yilun Du, Mengyu Wang
Abstract
As large language models (LLMs) transition from chat interfaces to integral components of stochastic pipelines and systems approaching general intelligence, the ability to faithfully sample from specified probability distributions has become a functional requirement rather than a theoretical curiosity. We present the first large-scale, statistically powered audit of native probabilistic sampling in frontier LLMs, benchmarking 11 models across 15 distributions. To disentangle failure modes, we employ a dual-protocol design: Batch Generation, where a model produces samples within one response, and Independent Requests, comprising stateless calls. We observe a sharp protocol asymmetry: batch generation achieves only modest statistical validity, with a 7% median pass rate, while independent requests collapse almost entirely, with 10 of 11 models passing none of the distributions. Beyond this asymmetry, we reveal that sampling fidelity degrades monotonically with distributional complexity and aggravates as the sampling horizon increases. Finally, we demonstrate how the propagation of these failures into downstream real-world application tasks introduces systematic biases: models fail to enforce uniform answer-position constraints in Multiple Choice Question generation and systematically violate demographic targets in attribute-constrained text-to-image prompt synthesis. These findings indicate that current LLMs lack a functional internal sampler, necessitating external tools for applications requiring statistical guarantees.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52c645a6-b515-491c-b507-1bf4dd4c2d3cBuilds on6
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming YinEMNLP 2023 · 102 citations
- Following Length Constraints in InstructionsWeizhe Yuan, Ilia Kulikov, Ping Yu, Kyunghyun Cho et al.EMNLP 2025 · 1 citation
- Large Language Models are not Fair EvaluatorsPeiyi Wang, Lei Li, Liang Chen, Zefan Cai et al.ACL 2024
Related papers
- Uncertainty Quantification for LLM-Based Survey SimulationsChengpiao Huang, Yuhang Wu, Kaizheng WangICML 2025
- Adaptive Generation of Bias-Eliciting Questions for LLMsRobin Staab, Jasper Dekoninck, Maximilian Baader, Martin VechevICML 2026
- D-Models and E-Models: Diversity-Stability Trade-offs in the Sampling Behavior of Large Language ModelsJia Gu, Liang Pang, Huawei Shen, Xueqi ChengWWW 2026
- CorrSynth - A Correlated Sampling Method for Diverse Dataset Generation from LLMsSuhas S. Kowshik, Abhishek Divekar, Vijit MalikEMNLP 2024
- Certifying Counterfactual Bias in LLMsIsha Chaudhary, Qian Hu, Manoj Kumar, Morteza Ziyadi et al.ICLR 2025 · 3 citations
