Is the Most Accurate AI the Best Teammate? Optimizing AI for Teamwork
Gagan Bansal, Besmira Nushi, Ece Kamar, Eric Horvitz, Daniel S. Weld
Abstract
AI practitioners typically strive to develop the most accurate systems, making an implicit assumption that the AI system will function autonomously. However, in practice, AI systems often are used to provide advice to people in domains ranging from criminal justice and finance to healthcare. In such AI-advised decision making, humans and machines form a team, where the human is responsible for making final decisions. But is the most accurate AI the best teammate? We argue "not necessarily" --- predictable performance may be worth a slight sacrifice in AI accuracy. Instead, we argue that AI systems should be trained in a human-centered manner, directly optimized for team performance. We study this proposal for a specific type of human-AI teaming, where the human overseer chooses to either accept the AI recommendation or solve the task themselves. To optimize the team performance for this setting we maximize the team's expected utility, expressed in terms of the quality of the final decision, cost of verifying, and individual accuracies of people and machines. Our experiments with linear and non-linear models on real-world, high-stakes datasets show that the most accuracy AI may not lead to highest team performance and show the benefit of modeling teamwork during training through improvements in expected team utility across datasets, considering parameters such as human skill and the cost of mistakes. We discuss the shortcoming of current optimization approaches beyond well-studied loss functions such as log-loss, and encourage future work on AI optimization problems motivated by human-AI collaboration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3084ba0-ecb9-4403-8a19-c5e0617f88b4Cited by top-tier papers32
- Two-Stage Learning to Defer with Multiple ExpertsAnqi Mao, Christopher Mohri, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 98 citations
- Combining Human Predictions with Model Probabilities via Confusion Matrices and CalibrationGavin Kerrigan, Padhraic Smyth, Mark SteyversNeurIPS 2021 · 79 citations
- Calibrated Learning to Defer with One-vs-All ClassifiersRajeev Verma, Eric T. NalisnickICML 2022 · 76 citations
- Post-hoc estimators for learning to defer to an expertHarikrishna Narasimhan, Wittawat Jitkrittum, Aditya Krishna Menon, Ankit Singh Rawat et al.NeurIPS 2022 · 66 citations
- Will You Accept the AI Recommendation? Predicting Human Behavior in AI-Assisted Decision MakingXinru Wang, Zhuoran Lu, Ming YinWWW 2022 · 58 citations
Related papers
- Uncalibrated Models Can Improve Human-AI CollaborationKailas Vodrahalli, Tobias Gerstenberg, James Y. ZouNeurIPS 2022 · 47 citations
- Align When They Want, Complement When They Need! Human-Centered Ensembles for Adaptive Human-AI CollaborationSyed Hasan Amin Mahmood, Ming Yin, Rajiv KhannaAAAI 2026
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- Utilizing Human Behavior Modeling to Manipulate Explanations in AI-Assisted Decision Making: The Good, the Bad, and the ScaryZhuoyan Li, Ming YinNeurIPS 2024 · 15 citations
- "Are You Really Sure?" Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision MakingShuai Ma, Xinru Wang, Ying Lei, Chuhan Shi et al.CHI 2024 · 54 citations
