To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language Models
Jessica Y. Bo, Sophia Wan, Ashton Anderson
摘要
As Large Language Models become integral to decision-making, optimism about their power is tempered with concern over their errors. Users may over-rely on LLM advice that is confidently stated but wrong, or under-rely due to mistrust. Reliance interventions have been developed to help users of LLMs, but they lack rigorous evaluation for appropriate reliance. We benchmark the performance of three relevant interventions by conducting a randomized online experiment with 400 participants attempting two challenging tasks: LSAT logical reasoning and image-based numerical estimation. For each question, participants first answered independently, then received LLM advice modified by one of three reliance interventions and answered the question again. Our findings indicate that while interventions reduce over-reliance, they generally fail to improve appropriate reliance. Furthermore, people became more confident after making wrong reliance decisions in certain contexts, demonstrating poor calibration. Based on our findings, we discuss implications for designing effective reliance interventions in human-LLM collaboration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving TasksJessica Y. Bo, Majeed Kazemitabaar, Mengqing Deng, Michael Inzlicht 等CHI 2026 · 被引用 8 次
- The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language ModelsYike Shi, Qing Xiao, Qing Hu, Hong Shen 等CHI 2026 · 被引用 7 次
- Behavioral Indicators of Overreliance During Interaction with Conversational Language ModelsChang Liu, Qinyi Zhou, Xinjie Shen, Xingyu Bruce Liu 等CHI 2026 · 被引用 4 次
- Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to ThemAllison Chen, Sunnie S. Y. Kim, Angel Nathaniel Franyutti-Cintron, Amaya Dharmasiri 等CHI 2026 · 被引用 3 次
- PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&AAnna Martin-Boyle, Cara A. C. Leckey, Martha Brown, Harmanpreet KaurCHI 2026 · 被引用 3 次
它引用的顶会 Paper30
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 被引用 962 次
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok 等CHI 2021 · 被引用 713 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Uncalibrated Models Can Improve Human-AI CollaborationKailas Vodrahalli, Tobias Gerstenberg, James Y. ZouNeurIPS 2022 · 被引用 47 次
- Belief Updating and Delegation in Multi-Task Human-AI Interaction: Evidence from Controlled SimulationsShreyan Biswas, Alexander Erlei, Ujwal GadirajuCHI 2026 · 被引用 4 次
- Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and InconsistenciesSunnie S. Y. Kim, Jennifer Wortman Vaughan, Q. Vera Liao, Tania Lombrozo 等CHI 2025 · 被引用 118 次
- Effects of LLM-based Search on Decision Making: Speed, Accuracy, and OverrelianceSofia Eleni Spatharioti, David M. Rothschild, Daniel G. Goldstein, Jake M. HofmanCHI 2025 · 被引用 28 次
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMsGabrielle Kaili-May Liu, Gal Yona, Avi Caciularu, Idan Szpektor 等EMNLP 2025
