Exploiting Opponents Under Utility Constraints in Sequential Games
Martino Bernasconi de Luca, Federico Cacciamani, Simone Fioravanti, Nicola Gatti, Alberto Marchesi, Francesco Trovò
摘要
Recently, game-playing agents based on AI techniques have demonstrated superhuman performance in several sequential games, such as chess, Go, and poker. Surprisingly, the multi-agent learning techniques that allowed to reach these achievements do not take into account the actual behavior of the human player, potentially leading to an impressive gap in performances. In this paper, we address the problem of designing artificial agents that learn how to effectively exploit unknown human opponents while playing repeatedly against them in an online fashion. We study the case in which the agent's strategy during each repetition of the game is subject to constraints ensuring that the human's expected utility is within some lower and upper thresholds. Our framework encompasses several real-world problems, such as human engagement in repeated game playing and human education by means of serious games. As a first result, we formalize a set of linear inequalities encoding the conditions that the agent's strategy must satisfy at each iteration in order to do not violate the given bounds for the human's expected utility. Then, we use such formulation in an upper confidence bound algorithm, and we prove that the resulting procedure suffers from sublinear regret and guarantees that the constraints are satisfied with high probability at each iteration. Finally, we empirically evaluate the convergence of our algorithm on standard testbeds of sequential games. * Equal contribution. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Sequential Information Design: Learning to Persuade in the DarkMartino Bernasconi, Matteo Castiglioni, Alberto Marchesi, Nicola Gatti 等NeurIPS 2022 · 被引用 19 次
- Safe Learning in Tree-Form Sequential Decision Making: Handling Hard and Soft ConstraintsMartino Bernasconi, Federico Cacciamani, Matteo Castiglioni, Alberto Marchesi 等ICML 2022 · 被引用 10 次
- Safe Opponent-Exploitation Subgame RefinementMingyang Liu, Chengjie Wu, Qihan Liu, Yansen Jing 等NeurIPS 2022 · 被引用 9 次
相关 Paper
- Learning to Play Sequential Games versus Unknown OpponentsPier Giuseppe Sessa, Ilija Bogunovic, Maryam Kamgarpour, Andreas KrauseNeurIPS 2020 · 被引用 34 次
- Online Learning in Unknown Markov GamesYi Tian, Yuanhao Wang, Tiancheng Yu, Suvrit SraICML 2021 · 被引用 48 次
- Human-Level Performance in No-Press Diplomacy via Equilibrium SearchJonathan Gray, Adam Lerer, Anton Bakhtin, Noam BrownICLR 2021 · 被引用 61 次
- Maximizing utility in multi-agent environments by anticipating the behavior of other learnersAngelos Assos, Yuval Dagan, Constantinos DaskalakisNeurIPS 2024 · 被引用 16 次
- Provable Self-Play Algorithms for Competitive Reinforcement LearningYu Bai, Chi JinICML 2020 · 被引用 169 次
