A Diachronic Perspective on User Trust in AI under Uncertainty
Shehzaad Dhuliawala, Vilém Zouhar, Mennatallah El-Assady, Mrinmaya Sachan
摘要
In human-AI collaboration, users typically form a mental model of the AI system, which captures the user’s beliefs about when the system performs well and when it does not. The construction of this mental model is guided by both the system’s veracity as well as the system output presented to the user e.g., the system’s confidence and an explanation for the prediction. However, modern NLP systems are seldom calibrated and are often confidently incorrect about their predictions, which violates users’ mental model and erodes their trust. In this work, we design a study where users bet on the correctness of an NLP system, and use it to study the evolution of user trust as a response to these trust-eroding events and how the user trust is rebuilt as a function of time after these events. We find that even a few highly inaccurate confidence estimation instances are enough to damage users’ trust in the system and performance, which does not easily recover over time. We further find that users are more forgiving to the NLP system if it is unconfidently correct rather than confidently incorrect, even though, from a game-theoretic perspective, their payoff is equivalent. Finally, we find that each user can entertain multiple mental models of the system based on the type of the question. These results highlight the importance of confidence calibration in developing user-centered NLP applications to avoid damaging user trust and compromising the collaboration performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- On the Robustness of Verbal Confidence of LLMs in Adversarial AttacksStephen Obadinma, Xiaodan ZhuNeurIPS 2025 · 被引用 3 次
- The Illusion of Competence: Evaluating the Effect of Explanations on Users' Mental Models of Visual Question Answering SystemsJudith Sieker, Simeon Junker, Ronja Utescher, Nazia Attari 等EMNLP 2024 · 被引用 2 次
- Rationalizing Transformer Predictions via End-To-End Differentiable Self-TrainingMarc Felix Brinner, Sina ZarrießEMNLP 2024
- Calibrating Large Language Models Using Their Generations OnlyDennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun 等ACL 2024
- Relying on the Unreliable: The Impact of Language Models' Reluctance to Express UncertaintyKaitlyn Zhou, Jena D. Hwang, Xiang Ren, Maarten SapACL 2024
它引用的顶会 Paper9
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok 等CHI 2021 · 被引用 713 次
- Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-MakingShuai Ma, Ying Lei, Xinru Wang, Chengbo Zheng 等CHI 2023 · 被引用 139 次
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 被引用 121 次
- When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning ModelsAmy Rechkemmer, Ming YinCHI 2022 · 被引用 94 次
- Teaching Humans When to Defer to a Classifier via ExemplarsHussein Mozannar, Arvind Satyanarayan, David A. SontagAAAI 2022 · 被引用 49 次
相关 Paper
- As Confidence Aligns: Understanding the Effect of AI Confidence on Human Self-confidence in Human-AI Decision MakingJingshu Li, Yitian Yang, Q. Vera Liao, Junti Zhang 等CHI 2025 · 被引用 56 次
- Uncalibrated Models Can Improve Human-AI CollaborationKailas Vodrahalli, Tobias Gerstenberg, James Y. ZouNeurIPS 2022 · 被引用 47 次
- To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language ModelsJessica Y. Bo, Sophia Wan, Ashton AndersonCHI 2025 · 被引用 31 次
- Evaluating What Others Say: The Effect of Accuracy Assessment in Shaping Mental Models of AI SystemsHyo Jin Do, Michelle Brachman, Casey Dugan, Qian Pan 等CSCW 2024 · 被引用 4 次
- Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving TasksJessica Y. Bo, Majeed Kazemitabaar, Mengqing Deng, Michael Inzlicht 等CHI 2026 · 被引用 8 次
