Explaining Sources of Uncertainty in Automated Fact-Checking
Jingyi Sun, Greta Warren, Irina Shklovski, Isabelle Augenstein
Abstract
Human-AI collaboration in knowledgeintensive tasks such as fact-checking requires understanding model uncertainty in multi-document reasoning amid conflicting/agreeing evidence. Yet, existing methods only express uncertainty as numbers or hedges without revealing which evidence conflicts cause the uncertainty, leaving users unable to resolve disagreements. We present CLUE (Conflict-&Agreement-aware Language-model Uncertainty Explanations), a plug-and-play white-box framework that, to our knowledge, is the first to generate natural-language explanations of model uncertainty grounded in conflicting/agreeing evidence. CLUE (i) identifies span-level claim-evidence and inter-evidence relations that signal conflict or agreement without supervision, and (ii) uses these relations to steer explanation generation, articulating how they drive the model's uncertainty. Across three language models and two fact-checking datasets, CLUE produces explanations that more faithfully track model uncertainty and better align with the model's fact-checking decisions than span-agnostic explanation prompting; human raters also judge them more helpful, more informative, less redundant, and more logically consistent with the input. By explicitly tying uncertainty to evidence conflicts and agreements, CLUE supports practical fact-checking and other tasks that require reasoning over complex, conflicting information. * Equal contribution. Claim: Scientific data has shown that cats can be infected with SARS-CoV-2 and can spread it to other cats. Model Output: Supports ✅ Model Certainty: 73% [...] there is a possibility of spreading SARS-CoV-2 through domestic pets Evidence 1 [...] no further transmission events to other animals or persons Evidence 2 Automated claim verification Span interactions for model uncertainty Claim: Scientific data has shown that cats can be infected with SARS-CoV-2 and can spread it to other cats.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08b2f668-0d5b-4961-978a-4db836daab7cCited by top-tier papers3
- They Think AI Can Do More Than It Actually Can: Practices, Challenges, & Opportunities of AI-Supported Reporting In Local JournalismBesjon Cifliku, Hendrik HeuerCHI 2026 · 2 citations
- Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development ResearchNimisha Karnatak, Mohamad Chatila, Daniel Alejandro Pinzón Hernández, Reza Yazdanfar et al.CHI 2026 · 1 citation
- A Reality Check on Context Utilisation for Retrieval-Augmented GenerationLovisa Hagström, Sara Vera Marjanovic, Haeun Yu, Arnav Arora et al.ACL 2025
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
Related papers
- Getting a CLUE: A Method for Explaining Uncertainty EstimatesJavier Antorán, Umang Bhatt, Tameem Adel, Adrian Weller et al.ICLR 2021 · 41 citations
- LM vs LM: Detecting Factual Errors via Cross ExaminationRoi Cohen, May Hamri, Mor Geva, Amir GlobersonEMNLP 2023 · 41 citations
- CER: Confidence Enhanced Reasoning in LLMsAli Razghandi, Seyed Mohammad Hadi Hosseini, Mahdieh Soleymani BaghshahACL 2025 · 11 citations
- A Fact-Checking Framework with Denoising Evidence Retrieval and LLM-Based Debate VerificationJun Yang, Yuhan Bai, Dandan Song, Zhijing Wu et al.WWW 2026
- Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial IntelligenceYanan Wang, Shuaicong Hu, Jian Liu, Guohui Zhou et al.ICML 2026
