Evaluation of Human-AI Teams for Learned and Rule-Based Agents in Hanabi
Ho Chit Siu, Jaime Daniel Peña, Edenna Chen, Yutai Zhou, Victor J. Lopez, Kyle Palko, Kimberlee C. Chang, Ross E. Allen
Abstract
Deep reinforcement learning has generated superhuman AI in competitive games such as Go and StarCraft. Can similar learning techniques create a superior AI teammate for human-machine collaborative games? Will humans prefer AI teammates that improve objective team performance or those that improve subjective metrics of trust? In this study, we perform a single-blind evaluation of teams of humans and AI agents in the cooperative card game Hanabi, with both rule-based and learning-based agents. In addition to the game score, used as an objective metric of the human-AI team performance, we also quantify subjective measures of the human's perceived performance, teamwork, interpretability, trust, and overall preference of AI teammate. We find that humans have a clear preference toward a rule-based AI teammate (SmartBot) over a state-of-the-art learning-based AI teammate (Other-Play) across nearly all subjective metrics, and generally view the learning-based agent negatively, despite no statistical difference in the game score. This result has implications for future AI design and reinforcement learning benchmarking, highlighting the need to incorporate subjective metrics of human-AI teaming rather than a singular focus on objective task performance. 4
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88297dae-7085-472d-8e2b-333a7d9ac540Cited by top-tier papers11
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Modeling Strong and Human-Like Gameplay with KL-Regularized SearchAthul Paul Jacob, David J. Wu, Gabriele Farina, Adam Lerer et al.ICML 2022 · 69 citations
- An Efficient End-to-End Training Approach for Zero-Shot Human-AI CoordinationXue Yan, Jiaxian Guo, Xingzhou Lou, Jun Wang et al.NeurIPS 2023 · 37 citations
- Diverse Conventions for Human-AI CollaborationBidipta Sarkar, Andy Shih, Dorsa SadighNeurIPS 2023 · 23 citations
- SimSpark: Interactive Simulation of Social Media BehaviorsZiyue Lin, Yi Shan, Lin Gao, Xinghua Jia et al.CSCW 2025 · 5 citations
Builds on5
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Simplified Action Decoder for Deep Multi-Agent Reinforcement LearningHengyuan Hu, Jakob N. FoersterICLR 2020 · 88 citations
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 87 citations
- Human-Level Performance in No-Press Diplomacy via Equilibrium SearchJonathan Gray, Adam Lerer, Anton Bakhtin, Noam BrownICLR 2021 · 61 citations
Related papers
- The Hidden Rules of Hanabi: How Humans Outperform AI AgentsMatthew Sidji, Wally Smith, Melissa J. RogersonCHI 2023 · 9 citations
- Learning to Lie: Adversarial Attacks on Human-AI Teams and LLMsAbed Kareem Musaffar, Anand Gokhale, Sirui Zeng, Rasta Tadayontahmasebi et al.ICLR 2026 · 2 citations
- Designs for Enabling Collaboration in Human-Machine Teaming via Interactive and Explainable SystemsRohan R. Paleja, Michael Munje, Kimberlee Chestnut Chang, Reed Jensen et al.NeurIPS 2024 · 12 citations
- K-level Reasoning for Zero-Shot Coordination in HanabiBrandon Cui, Hengyuan Hu, Luis Pineda, Jakob N. FoersterNeurIPS 2021 · 46 citations
- Investigating AI Teammate Communication Strategies and Their Impact in Human-AI Teams for Effective TeamworkRui Zhang, Wen Duan, Christopher Flathmann, Nathan J. McNeese et al.CSCW 2023 · 98 citations
