Perceptions of Linguistic Uncertainty by Language Models and Humans
Catarina G. Belém, Markelle Kelly, Mark Steyvers, Sameer Singh, Padhraic Smyth
摘要
Uncertainty expressions such as "probably" or "highly unlikely" are pervasive in human language. While prior work has established that there is population-level agreement in terms of how humans quantitatively interpret these expressions, there has been little inquiry into the abilities of language models in the same context. In this paper, we investigate how language models map linguistic expressions of uncertainty to numerical responses. Our approach assesses whether language models can employ theory of mind in this setting: understanding the uncertainty of another agent about a particular statement, independently of the model's own certainty about that statement. We find that 7 out of 10 models are able to map uncertainty expressions to probabilistic responses in a human-like manner. However, we observe systematically different behavior depending on whether a statement is actually true or false. This sensitivity indicates that language models are substantially more susceptible to bias based on their prior knowledge (as compared to humans). These findings raise important questions and have broad implications for human-AI and AI-AI communication.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Logical forms complement probability in understanding language model (and human) performanceYixuan Wang, Freda ShiACL 2025 · 被引用 2 次
- Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and AttitudesMeng Li, Michael Vrazitulis, David SchlangenACL 2025 · 被引用 1 次
- “very likely” Means “uncertain”? How LLMs Diverge from Humans in Linguistic Uncertainty QuantificationJinhao Duan, Zicheng Liu, Zijie Liu, Kaidi Xu 等ICML 2026
- Demystifying Uncertainty in LLMs: Active Calibration between Concepts and Human EvaluationsPengqi Li, Lizhong Ding, Zhehao Zhou, Chunhui Zhang 等ACL 2026
它引用的顶会 Paper13
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 被引用 651 次
相关 Paper
- Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others from Conversational CuesAnthony B. Sicilia, Malihe AlikhaniACL 2025 · 被引用 2 次
- Calibrating Expressions of CertaintyPeiqi Wang, Barbara D. Lam, Yingcheng Liu, Ameneh Asgari-Targhi 等ICLR 2025
- Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language ModelsKaitlyn Zhou, Dan Jurafsky, Tatsunori HashimotoEMNLP 2023 · 被引用 29 次
- Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMsSean Trott, Samuel M. Taylor, Cameron Robert Jones, James A. Michaelov 等ACL 2026 · 被引用 2 次
- SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?Michael Kirchhof, Luca Füger, Adam Golinski, Eeshan Gunesh Dhekane 等ICLR 2026 · 被引用 4 次
