Language Model Probabilities are Not Calibrated in Numeric Contexts
Charles Lovering, Michael Krumdick, Viet Dac Lai, Varshini Reddy, Seth Ebner, Nilesh Kumar, Rik Koncel-Kedziorski, Chris Tanner
Abstract
Some statements have one well-defined continuation (e.g., "the Eiffel Tower is in [Paris]"), whereas others have a natural distribution over multiple options (e.g., "the weighted coin flip was [Heads/Tails].") We argue that language model (LM) outputs should capture these natural distributions. Our work specifically tests whether LM output probabilities are calibrated to numeric information within their textual contexts. For example, if the context (the prompt) concerns two equally likely options (e.g., heads or tails for a fair coin), the LM output probabilities should also be equal. Likewise, in a context with nonuniformly likely events (e.g., rolling a pair with two dice) an LM should output proportionate probabilities. However, we find that even in simple settings, the best LMs (1) are poorly calibrated and (2) have systematic biases: artifacts like word identity, word order, and word frequency all impact calibration. For example, gpt-4o-mini often picks the first of two options presented in the prompt regardless of the options' implied likelihoods, whereas Llama-3.1-8B picks the second. Models do not allocate probability mass among valid options in a calibrated manner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 482dcfc4-0438-42b9-84d4-8ec2f62d33ddCited by top-tier papers3
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse GenerationKou Misaki, Takuya AkibaICLR 2026 · 13 citations
- Frontier Models Can Take Actions at Low ProbabilitiesAlex Serrano Terre, Wen Xing, David Lindner, Erik JennerICML 2026 · 1 citation
- Entropy-informed Decoding: Adaptive Information-Driven BranchingBenjamin Patrick Evans, Sumitra Ganesh, Leo ArdonICML 2026
Related papers
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Linguistic Calibration of Long-Form GenerationsNeil Band, Xuechen Li, Tengyu Ma, Tatsunori HashimotoICML 2024 · 56 citations
- Are Language Models Any Good at Density Modeling?Sriram Ranga, Sai Shashank Bedampeta, Rui Mao, Anupam ChattopadhyayAAAI 2026
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 2 citations
- Evaluating Distributional Distortion in Neural Language ModelingBenjamin LeBrun, Alessandro Sordoni, Timothy J. O'DonnellICLR 2022 · 26 citations
