Too Sure for Our Own Good: A User Study on AI Confidence and Human Reliance
Caterina Fregosi, Lucia Vicente, Andrea Campagner, Federico Cabitza
Abstract
Achieving appropriate human reliance on Artificial Intelligence (AI) systems remains a central challenge in Human-Computer Interaction. Confidence scores—indicators of an AI system’s certainty in its recommendations—have been proposed as a means to help users calibrate their trust and reliance on AI Decision Support Systems (DSS). However, limited research has explored how well-calibrated versus miscalibrated confidence scores affect human decision-making. We report a study examining the effects of confidence calibration on user reliance, decision accuracy, and perceived utility of an AI DSS. In a within-subjects experiment involving 184 participants solving logic puzzles, we found that well-calibrated confidence scores significantly improved decision accuracy (+20%, 95% CI: [0.18, 0.23]), whereas miscalibrated scores yielded minimal accuracy gains (+2%, 95% CI: [-0.00, 0.04]) and increased vulnerability to automation bias and conservatism bias. Participants were more likely to accept AI recommendations when high confidence was expressed, even when those recommendations were incorrect, resulting in errors. Conversely, miscalibrated and low-confidence recommendations increased conservatism bias, leading users to reject even accurate AI suggestions. Perceived utility of the AI system was higher when confidence levels were high (p < 0.001) and when confidence was well-calibrated (p = 0.002). These findings underscore the importance of designing AI systems with properly calibrated confidence cues to improve human-AI collaboration and mitigate reliance-related biases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f254c5f-bbaf-432f-93b8-38bbccdeb7b6Builds on6
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- Explanations Can Reduce Overreliance on AI Systems During Decision-MakingHelena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg et al.CSCW 2023 · 362 citations
- Explanations, Fairness, and Appropriate Reliance in Human-AI Decision-MakingJakob Schoeffer, Maria De-Arteaga, Niklas KühlCHI 2024 · 67 citations
- Uncalibrated Models Can Improve Human-AI CollaborationKailas Vodrahalli, Tobias Gerstenberg, James Y. ZouNeurIPS 2022 · 47 citations
- Designing for Appropriate Reliance: The Roles of AI Uncertainty Presentation, Initial User Decision, and User Demographics in AI-Assisted Decision-MakingShiye Cao, Anqi Liu, Chien-Ming HuangCSCW 2024 · 38 citations
Related papers
- "Are You Really Sure?" Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision MakingShuai Ma, Xinru Wang, Ying Lei, Chuhan Shi et al.CHI 2024 · 54 citations
- As Confidence Aligns: Understanding the Effect of AI Confidence on Human Self-confidence in Human-AI Decision MakingJingshu Li, Yitian Yang, Q. Vera Liao, Junti Zhang et al.CHI 2025 · 56 citations
- Knowing About Knowing: An Illusion of Human Competence Can Hinder Appropriate Reliance on AI SystemsGaole He, Lucie Kuiper, Ujwal GadirajuCHI 2023 · 101 citations
- To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language ModelsJessica Y. Bo, Sophia Wan, Ashton AndersonCHI 2025 · 31 citations
- Human-Aligned Calibration for AI-Assisted Decision MakingNina Corvelo Benz, Manuel Gomez RodriguezNeurIPS 2023 · 45 citations
