Measuring Sentence-Level and Aspect-Level (Un)certainty in Science Communications
Jiaxin Pei, David Jurgens
Abstract
Certainty and uncertainty are fundamental to science communication. Hedges have widely been used as proxies for uncertainty. However, certainty is a complex construct, with authors expressing not only the degree but the type and aspects of uncertainty in order to give the reader a certain impression of what is known. Here, we introduce a new study of certainty that models both the level and the aspects of certainty in scientific findings. Using a new dataset of 2167 annotated scientific findings, we demonstrate that hedges alone account for only a partial explanation of certainty. We show that both the overall certainty and individual aspects can be predicted with pre-trained language models, providing a more complete picture of the author's intended communication. Downstream analyses on 431K scientific findings from news and scientific abstracts demonstrate that modeling sentencelevel and aspect-level certainty is meaningful for areas like science communication. Both the model and datasets used in this paper are released at https://blablablab.si. umich.edu/projects/certainty/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6fadeb0d-e9e6-428d-b522-0f28a4992fe8Cited by top-tier papers6
- Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET BenchmarkMinje Choi, Jiaxin Pei, Sagar Kumar, Chang Shu et al.EMNLP 2023 · 36 citations
- Inference to the Best Explanation in Large Language ModelsDhairya Dalal, Marco Valentino, André Freitas, Paul BuitelaarACL 2024 · 2 citations
- Modeling Information Change in Science Communication with Semantically Matched ParaphrasesDustin Wright, Jiaxin Pei, David Jurgens, Isabelle AugensteinEMNLP 2022 · 2 citations
- Counterfactual LLM-based Framework for Measuring Rhetorical StyleJingyi Qiu, Hong Chen, Zongyi LiICLR 2026 · 1 citation
- Missci: Reconstructing Fallacies in Misrepresented ScienceMax Glockner, Yufang Hou, Preslav Nakov, Iryna GurevychACL 2024 · 1 citation
Related papers
- Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language ModelsKaitlyn Zhou, Dan Jurafsky, Tatsunori HashimotoEMNLP 2023 · 29 citations
- SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?Michael Kirchhof, Luca Füger, Adam Golinski, Eeshan Gunesh Dhekane et al.ICLR 2026 · 4 citations
- Perceptions of Linguistic Uncertainty by Language Models and HumansCatarina G. Belém, Markelle Kelly, Mark Steyvers, Sameer Singh et al.EMNLP 2024 · 6 citations
- Calibrating Expressions of CertaintyPeiqi Wang, Barbara D. Lam, Yingcheng Liu, Ameneh Asgari-Targhi et al.ICLR 2025
- Distinguishing the Knowable from the Unknowable with Language ModelsGustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak et al.ICML 2024 · 44 citations
