Learning to Lie: Adversarial Attacks on Human-AI Teams and LLMs
Abed Kareem Musaffar, Anand Gokhale, Sirui Zeng, Rasta Tadayontahmasebi, Xifeng Yan, Ambuj K. Singh, Francesco Bullo
摘要
As artificial intelligence (AI) assistants become more widely adopted in safety-critical domains, it becomes important to develop safeguards against potential failures or adversarial attacks. A key prerequisite to developing these safeguards is understanding the ability of these AI assistants to mislead human teammates. We investigate this attack problem within the context of an intellective strategy game where a team of three humans and one AI assistant collaborate to answer a series of trivia questions. Unbeknownst to the humans, the AI assistant is adversarial. Leveraging techniques from Model-Based Reinforcement Learning (MBRL), the AI assistant learns a model of the humans' trust evolution and uses that model to manipulate the group decision-making process to harm the team. We evaluate two models---one inspired by literature and the other data-driven---and find that both can effectively harm the human team. Moreover, we find that in this setting while our data-driven model is the most capable of accurately predicting how human agents appraise their teammates given limited information on prior interactions, the model based on principles of cognitive psychology does not lag too far behind. Finally, we compare the performance of state-of-the-art LLM models to human agents on our influence allocation task to evaluate whether the LLMs allocate influence similarly to humans or if they are more robust to our attack. These results enhance our understanding of decision-making dynamics in small human-AI teams and lay the foundation for defense strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Deciding Fast and Slow: The Role of Cognitive Biases in AI-assisted Decision-makingCharvi Rastogi, Yunfeng Zhang, Dennis Wei, Kush R. Varshney 等CSCW 2022 · 被引用 184 次
- Fooling Detection Alone is Not Enough: Adversarial Attack against Multiple Object TrackingYunhan Jia, Yantao Lu, Junjie Shen, Qi Alfred Chen 等ICLR 2020 · 被引用 113 次
- Modeling Human Trust and Reliance in AI-Assisted Decision Making: A Markovian ApproachZhuoyan Li, Zhuoran Lu, Ming YinAAAI 2023 · 被引用 28 次
相关 Paper
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS 等ICML 2026 · 被引用 4 次
- Teaching Humans When to Defer to a Classifier via ExemplarsHussein Mozannar, Arvind Satyanarayan, David A. SontagAAAI 2022 · 被引用 49 次
- Measuring and Mitigating Rapport Bias of Large Language Models under Multi-Agent Social InteractionsMaojia Song, Pala Tej Deep, Ruiwen Zhou, Weisheng Jin 等ICLR 2026
- Evaluation of Human-AI Teams for Learned and Rule-Based Agents in HanabiHo Chit Siu, Jaime Daniel Peña, Edenna Chen, Yutai Zhou 等NeurIPS 2021 · 被引用 78 次
- Investigating AI Teammate Communication Strategies and Their Impact in Human-AI Teams for Effective TeamworkRui Zhang, Wen Duan, Christopher Flathmann, Nathan J. McNeese 等CSCW 2023 · 被引用 98 次
