Safe Opponent-Exploitation Subgame Refinement
Mingyang Liu, Chengjie Wu, Qihan Liu, Yansen Jing, Jun Yang, Pingzhong Tang, Chongjie Zhang
摘要
In zero-sum games, an NE strategy tends to be overly conservative confronted with opponents of limited rationality, because it does not actively exploit their weaknesses. From another perspective, best responding to an estimated opponent model is vulnerable to estimation errors and lacks safety guarantees. Inspired by the recent success of real-time search algorithms in developing superhuman AI, we investigate the dilemma of safety and opponent exploitation and present a novel real-time search framework, called Safe Exploitation Search (SES), which continuously interpolates between the two extremes of online strategy refinement. We provide SES with a theoretically upper-bounded exploitability and a lowerbounded evaluation performance. Additionally, SES enables computationally efficient online adaptation to a possibly updating opponent model, while previous safe exploitation methods have to recompute for the whole game. Empirical results show that SES significantly outperforms NE baselines and previous algorithms while keeping exploitability low at the same time. * Equal contribution 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Conservative Offline Policy Adaptation in Multi-Agent GamesChengjie Wu, Pingzhong Tang, Jun Yang, Yujing Hu 等NeurIPS 2023 · 被引用 4 次
- Safe and Robust Subgame Exploitation in Imperfect Information GamesZhenxing Ge, Zheng Xu, Tianyu Ding, Linjian Meng 等ICML 2024 · 被引用 3 次
- Best of Both Worlds: Regret Minimization versus Minimax PlayAdrian Müller, Jon Schneider, Stratis Skoulakis, Luca Viano 等ICML 2025
它引用的顶会 Paper4
- Combining Deep Reinforcement Learning and Search for Imperfect-Information GamesNoam Brown, Anton Bakhtin, Adam Lerer, Qucheng GongNeurIPS 2020 · 被引用 205 次
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 被引用 87 次
- Joint Policy Search for Multi-agent Collaboration with Imperfect InformationYuandong Tian, Qucheng Gong, Yu JiangNeurIPS 2020 · 被引用 24 次
- Exploiting Opponents Under Utility Constraints in Sequential GamesMartino Bernasconi de Luca, Federico Cacciamani, Simone Fioravanti, Nicola Gatti 等NeurIPS 2021 · 被引用 10 次
相关 Paper
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS 等ICML 2026 · 被引用 4 次
- Trajectory-wise Iterative Reinforcement Learning Framework for Auto-biddingHaoming Li, Yusen Huo, Shuai Dou, Zhenzhe Zheng 等WWW 2024 · 被引用 11 次
- Safe Learning in Tree-Form Sequential Decision Making: Handling Hard and Soft ConstraintsMartino Bernasconi, Federico Cacciamani, Matteo Castiglioni, Alberto Marchesi 等ICML 2022 · 被引用 10 次
- Maximize to Explore: One Objective Function Fusing Estimation, Planning, and ExplorationZhihan Liu, Miao Lu, Wei Xiong, Han Zhong 等NeurIPS 2023 · 被引用 30 次
- Safe Exploration via Policy PriorsManuel Wendl, Yarden As, Manish Prajapat, Anton Pollak 等ICLR 2026 · 被引用 6 次
