Enough Coin Flips Can Make LLMs Act Bayesian
Ritwik Gupta, Rodolfo Corona, Jiaxin Ge, Eric Wang, Dan Klein, Trevor Darrell, David M. Chan
Abstract
Large language models (LLMs) exhibit the ability to generalize given few-shot examples in their input prompt, an emergent capability known as in-context learning (ICL). We investigate whether LLMs use ICL to perform structured reasoning in ways that are consistent with a Bayesian framework or rely on pattern matching. Using a controlled setting of biased coin flips, we find that: (1) LLMs often possess biased priors, causing initial divergence in zero-shot settings, (2) in-context evidence outweighs explicit bias instructions, (3) LLMs broadly follow Bayesian posterior updates, with deviations primarily due to miscalibrated priors rather than flawed updates, and (4) attention magnitude has negligible effect on Bayesian inference. With sufficient demonstrations of biased coin flips via ICL, LLMs update their priors in a Bayesian manner. Code and visualizations are available on the project page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4944476-d01f-4662-8a30-082cc4b6ff04Cited by top-tier papers8
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM DiversityJiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia et al.ICML 2026 · 102 citations
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse GenerationKou Misaki, Takuya AkibaICLR 2026 · 13 citations
- Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM ReasoningZhonghao He, Tianyi Alex Qiu, Hirokazu Shirado, Maarten SapNeurIPS 2025 · 6 citations
- Pre-trained Large Language Models Learn to Predict Hidden Markov Models In-contextYijia Dai, Zhaolin Gao, Yahya Sattar, Sarah Dean et al.NeurIPS 2025 · 2 citations
- D-Models and E-Models: Diversity-Stability Trade-offs in the Sampling Behavior of Large Language ModelsJia Gu, Liang Pang, Huawei Shen, Xueqi ChengWWW 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
Related papers
- Many-Shot In-Context LearningRishabh Agarwal, Avi Singh, Lei Zhang, Bernd Bohnet et al.NeurIPS 2024 · 271 citations
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation SteeringEric Bigelow, Daniel Wurgaft, YingQiao Wang, Noah Goodman et al.ICML 2026
- Universal Self-Adaptive PromptingXingchen Wan, Ruoxi Sun, Hootan Nakhost, Hanjun Dai et al.EMNLP 2023 · 4 citations
- Measuring Inductive Biases of In-Context Learning with Underspecified DemonstrationsChenglei Si, Dan Friedman, Nitish Joshi, Shi Feng et al.ACL 2023 · 8 citations
- What Do Language Models Learn in Context? The Structured Task HypothesisJiaoda Li, Yifan Hou, Mrinmaya Sachan, Ryan CotterellACL 2024 · 5 citations
