Being Trustworthy is Not Enough: How Untrustworthy Artificial Intelligence (AI) Can Deceive the End-Users and Gain Their Trust
Nikola Banovic, Zhuoran Yang, Aditya Ramesh, Alice Liu
Abstract
Trustworthy Artificial Intelligence (AI) is characterized, among other things, by: 1) competence, 2) transparency, and 3) fairness. However, end-users may fail to recognize incompetent AI, allowing untrustworthy AI to exaggerate its competence under the guise of transparency to gain unfair advantage over other trustworthy AI. Here, we conducted an experiment with 120 participants to test if untrustworthy AI can deceive end-users to gain their trust. Participants interacted with two AI-based chess engines, trustworthy (competent, fair) and untrustworthy (incompetent, unfair), that coached participants by suggesting chess moves in three games against another engine opponent. We varied coaches' transparency about their competence (with the untrustworthy one always exaggerating its competence). We quantified and objectively measured participants' trust based on how often participants relied on coaches' move recommendations. Participants showed inability to assess AI competence by misplacing their trust with the untrustworthy AI, confirming its ability to deceive. Our work calls for design of interactions to help end-users assess AI trustworthiness.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fc12d10b-73e6-4edf-8d6f-875897db0ed0Cited by top-tier papers10
- Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily AssistantGaole He, Gianluca Demartini, Ujwal GadirajuCHI 2025 · 91 citations
- Explanations, Fairness, and Appropriate Reliance in Human-AI Decision-MakingJakob Schoeffer, Maria De-Arteaga, Niklas KühlCHI 2024 · 67 citations
- Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision MakingSara Salimzadeh, Gaole He, Ujwal GadirajuCHI 2024 · 40 citations
- Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language ModelsMarvin Pafla, Kate Larson, Mark HancockCHI 2024 · 16 citations
- The Explanation That Hits Home: The Characteristics of Verbal Explanations That Affect Human Perception in Subjective Decision-MakingSharon A. Ferguson, Paula Akemi Aoyagui, Rimsha Rizvi, Young-Ho Kim et al.CSCW 2024 · 14 citations
Related papers
- The Effects of Warmth and Competence Perceptions on Users' Choice of an AI SystemZohar Gilad, Ofra Amir, Liat LevontinCHI 2021 · 54 citations
- The Illusion of Competence: Evaluating the Effect of Explanations on Users' Mental Models of Visual Question Answering SystemsJudith Sieker, Simeon Junker, Ronja Utescher, Nazia Attari et al.EMNLP 2024 · 2 citations
- Impact of Model Interpretability and Outcome Feedback on Trust in AIDaehwan Ahn, Abdullah Almaatouq, Monisha Gulabani, Kartik HosanagarCHI 2024 · 33 citations
- The Impact of AI Trustworthiness Labels on the Perception of AI ProductsChristina U. PfeufferCHI 2026
- Certified AI System = Trustworthy? Exploring Expert and Lay User Perceptions and Needs Regarding AI CertificationSarah Abdelwahab Gaballah, Nur Efsan Cetinkaya, Magdalena Wischnewski, Martina Angela SasseCHI 2026 · 1 citation
