Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models
Kaitlyn Zhou, Dan Jurafsky, Tatsunori Hashimoto
Abstract
The increased deployment of LMs for real-world tasks involving knowledge and facts makes it important to understand model epistemology: what LMs think they know, and how their attitudes toward that knowledge are affected by language use in their inputs. Here, we study an aspect of model epistemology: how epistemic markers of certainty, uncertainty, or evidentiality like "I'm sure it's", "I think it's", or "Wikipedia says it's" affect models, and whether they contribute to model failures. We develop a typology of epistemic markers and inject 50 markers into prompts for question answering. We find that LMs are highly sensitive to epistemic markers in prompts, with accuracies varying more than 80%. Surprisingly, we find that expressions of high certainty result in a 7% decrease in accuracy as compared to low certainty expressions; similarly, factive verbs hurt performance, while evidentials benefit performance. Our analysis of a popular pretraining dataset shows that these markers of uncertainty are associated with answers on question-answering websites, while markers of certainty are associated with questions. These associations may suggest that the behavior of LMs is based on mimicking observed language use, rather than truly reflecting epistemic uncertainty.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a96d78e-b8b6-450d-ae82-85fb09d1ae43Cited by top-tier papers35
- Linguistic Calibration of Long-Form GenerationsNeil Band, Xuechen Li, Tengyu Ma, Tatsunori HashimotoICML 2024 · 56 citations
- Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM CollaborationShangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding et al.ACL 2024 · 30 citations
- LACIE: Listener-Aware Finetuning for Calibration in Large Language ModelsElias Stengel-Eskin, Peter Hase, Mohit BansalNeurIPS 2024 · 26 citations
- Robust Hallucination Detection in LLMs via Adaptive Token SelectionMengjia Niu, Hamed Haddadi, Guansong PangNeurIPS 2025 · 24 citations
- SteerConf: Steering LLMs for Confidence ElicitationZiang Zhou, Tianyuan Jin, Jieming Shi, Qing LiNeurIPS 2025 · 23 citations
Builds on12
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis et al.NeurIPS 2021 · 633 citations
- Social Simulacra: Creating Populated Prototypes for Social Computing SystemsJoon Sung Park, Lindsay Popowski, Carrie J. Cai, Meredith Ringel Morris et al.UIST 2022 · 192 citations
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 121 citations
- On the Inference Calibration of Neural Machine TranslationShuo Wang, Zhaopeng Tu, Shuming Shi, Yang LiuACL 2020 · 66 citations
Related papers
- Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and AttitudesMeng Li, Michael Vrazitulis, David SchlangenACL 2025 · 1 citation
- It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of BeliefKevin Du, Clara Kümpel, Michelle Wastl, Alex WarstadtACL 2026
- Perceptions of Linguistic Uncertainty by Language Models and HumansCatarina G. Belém, Markelle Kelly, Mark Steyvers, Sameer Singh et al.EMNLP 2024 · 6 citations
- To Believe or Not to Believe Your LLM: Iterative Prompting for Estimating Epistemic UncertaintyYasin Abbasi-Yadkori, Ilja Kuzborskij, András György, Csaba SzepesváriNeurIPS 2024
- CertainlyUncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric AwarenessKhyathi Raghavi Chandu, Linjie Li, Anas Awadalla, Ximing Lu et al.ICLR 2025
