To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language Models
Jessica Y. Bo, Sophia Wan, Ashton Anderson
Abstract
As Large Language Models become integral to decision-making, optimism about their power is tempered with concern over their errors. Users may over-rely on LLM advice that is confidently stated but wrong, or under-rely due to mistrust. Reliance interventions have been developed to help users of LLMs, but they lack rigorous evaluation for appropriate reliance. We benchmark the performance of three relevant interventions by conducting a randomized online experiment with 400 participants attempting two challenging tasks: LSAT logical reasoning and image-based numerical estimation. For each question, participants first answered independently, then received LLM advice modified by one of three reliance interventions and answered the question again. Our findings indicate that while interventions reduce over-reliance, they generally fail to improve appropriate reliance. Furthermore, people became more confident after making wrong reliance decisions in certain contexts, demonstrating poor calibration. Based on our findings, we discuss implications for designing effective reliance interventions in human-LLM collaboration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e10973b5-e906-40dd-90ef-2a8895e9c646Cited by top-tier papers10
- Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving TasksJessica Y. Bo, Majeed Kazemitabaar, Mengqing Deng, Michael Inzlicht et al.CHI 2026 · 8 citations
- The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language ModelsYike Shi, Qing Xiao, Qing Hu, Hong Shen et al.CHI 2026 · 7 citations
- Behavioral Indicators of Overreliance During Interaction with Conversational Language ModelsChang Liu, Qinyi Zhou, Xinjie Shen, Xingyu Bruce Liu et al.CHI 2026 · 4 citations
- Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to ThemAllison Chen, Sunnie S. Y. Kim, Angel Nathaniel Franyutti-Cintron, Amaya Dharmasiri et al.CHI 2026 · 3 citations
- PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&AAnna Martin-Boyle, Cara A. C. Leckey, Martha Brown, Harmanpreet KaurCHI 2026 · 3 citations
Builds on30
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 962 citations
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Uncalibrated Models Can Improve Human-AI CollaborationKailas Vodrahalli, Tobias Gerstenberg, James Y. ZouNeurIPS 2022 · 47 citations
- Belief Updating and Delegation in Multi-Task Human-AI Interaction: Evidence from Controlled SimulationsShreyan Biswas, Alexander Erlei, Ujwal GadirajuCHI 2026 · 4 citations
- Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and InconsistenciesSunnie S. Y. Kim, Jennifer Wortman Vaughan, Q. Vera Liao, Tania Lombrozo et al.CHI 2025 · 118 citations
- Effects of LLM-based Search on Decision Making: Speed, Accuracy, and OverrelianceSofia Eleni Spatharioti, David M. Rothschild, Daniel G. Goldstein, Jake M. HofmanCHI 2025 · 28 citations
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMsGabrielle Kaili-May Liu, Gal Yona, Avi Caciularu, Idan Szpektor et al.EMNLP 2025
