Learning to Lie: Adversarial Attacks on Human-AI Teams and LLMs
Abed Kareem Musaffar, Anand Gokhale, Sirui Zeng, Rasta Tadayontahmasebi, Xifeng Yan, Ambuj K. Singh, Francesco Bullo
Abstract
As artificial intelligence (AI) assistants become more widely adopted in safety-critical domains, it becomes important to develop safeguards against potential failures or adversarial attacks. A key prerequisite to developing these safeguards is understanding the ability of these AI assistants to mislead human teammates. We investigate this attack problem within the context of an intellective strategy game where a team of three humans and one AI assistant collaborate to answer a series of trivia questions. Unbeknownst to the humans, the AI assistant is adversarial. Leveraging techniques from Model-Based Reinforcement Learning (MBRL), the AI assistant learns a model of the humans' trust evolution and uses that model to manipulate the group decision-making process to harm the team. We evaluate two models---one inspired by literature and the other data-driven---and find that both can effectively harm the human team. Moreover, we find that in this setting while our data-driven model is the most capable of accurately predicting how human agents appraise their teammates given limited information on prior interactions, the model based on principles of cognitive psychology does not lag too far behind. Finally, we compare the performance of state-of-the-art LLM models to human agents on our influence allocation task to evaluate whether the LLMs allocate influence similarly to humans or if they are more robust to our attack. These results enhance our understanding of decision-making dynamics in small human-AI teams and lay the foundation for defense strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b59b8d30-3da6-4583-a4de-ee19ad3cb5dfBuilds on3
- Deciding Fast and Slow: The Role of Cognitive Biases in AI-assisted Decision-makingCharvi Rastogi, Yunfeng Zhang, Dennis Wei, Kush R. Varshney et al.CSCW 2022 · 184 citations
- Fooling Detection Alone is Not Enough: Adversarial Attack against Multiple Object TrackingYunhan Jia, Yantao Lu, Junjie Shen, Qi Alfred Chen et al.ICLR 2020 · 113 citations
- Modeling Human Trust and Reliance in AI-Assisted Decision Making: A Markovian ApproachZhuoyan Li, Zhuoran Lu, Ming YinAAAI 2023 · 28 citations
Related papers
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS et al.ICML 2026 · 4 citations
- Teaching Humans When to Defer to a Classifier via ExemplarsHussein Mozannar, Arvind Satyanarayan, David A. SontagAAAI 2022 · 49 citations
- Measuring and Mitigating Rapport Bias of Large Language Models under Multi-Agent Social InteractionsMaojia Song, Pala Tej Deep, Ruiwen Zhou, Weisheng Jin et al.ICLR 2026
- Evaluation of Human-AI Teams for Learned and Rule-Based Agents in HanabiHo Chit Siu, Jaime Daniel Peña, Edenna Chen, Yutai Zhou et al.NeurIPS 2021 · 78 citations
- Investigating AI Teammate Communication Strategies and Their Impact in Human-AI Teams for Effective TeamworkRui Zhang, Wen Duan, Christopher Flathmann, Nathan J. McNeese et al.CSCW 2023 · 98 citations
