Safe Opponent-Exploitation Subgame Refinement
Mingyang Liu, Chengjie Wu, Qihan Liu, Yansen Jing, Jun Yang, Pingzhong Tang, Chongjie Zhang
Abstract
In zero-sum games, an NE strategy tends to be overly conservative confronted with opponents of limited rationality, because it does not actively exploit their weaknesses. From another perspective, best responding to an estimated opponent model is vulnerable to estimation errors and lacks safety guarantees. Inspired by the recent success of real-time search algorithms in developing superhuman AI, we investigate the dilemma of safety and opponent exploitation and present a novel real-time search framework, called Safe Exploitation Search (SES), which continuously interpolates between the two extremes of online strategy refinement. We provide SES with a theoretically upper-bounded exploitability and a lowerbounded evaluation performance. Additionally, SES enables computationally efficient online adaptation to a possibly updating opponent model, while previous safe exploitation methods have to recompute for the whole game. Empirical results show that SES significantly outperforms NE baselines and previous algorithms while keeping exploitability low at the same time. * Equal contribution 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Conservative Offline Policy Adaptation in Multi-Agent GamesChengjie Wu, Pingzhong Tang, Jun Yang, Yujing Hu et al.NeurIPS 2023 · 4 citations
- Safe and Robust Subgame Exploitation in Imperfect Information GamesZhenxing Ge, Zheng Xu, Tianyu Ding, Linjian Meng et al.ICML 2024 · 3 citations
- Best of Both Worlds: Regret Minimization versus Minimax PlayAdrian Müller, Jon Schneider, Stratis Skoulakis, Luca Viano et al.ICML 2025
Builds on4
- Combining Deep Reinforcement Learning and Search for Imperfect-Information GamesNoam Brown, Anton Bakhtin, Adam Lerer, Qucheng GongNeurIPS 2020 · 205 citations
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 87 citations
- Joint Policy Search for Multi-agent Collaboration with Imperfect InformationYuandong Tian, Qucheng Gong, Yu JiangNeurIPS 2020 · 24 citations
- Exploiting Opponents Under Utility Constraints in Sequential GamesMartino Bernasconi de Luca, Federico Cacciamani, Simone Fioravanti, Nicola Gatti et al.NeurIPS 2021 · 10 citations
Related papers
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS et al.ICML 2026 · 4 citations
- Trajectory-wise Iterative Reinforcement Learning Framework for Auto-biddingHaoming Li, Yusen Huo, Shuai Dou, Zhenzhe Zheng et al.WWW 2024 · 11 citations
- Safe Learning in Tree-Form Sequential Decision Making: Handling Hard and Soft ConstraintsMartino Bernasconi, Federico Cacciamani, Matteo Castiglioni, Alberto Marchesi et al.ICML 2022 · 10 citations
- Maximize to Explore: One Objective Function Fusing Estimation, Planning, and ExplorationZhihan Liu, Miao Lu, Wei Xiong, Han Zhong et al.NeurIPS 2023 · 30 citations
- Safe Exploration via Policy PriorsManuel Wendl, Yarden As, Manish Prajapat, Anton Pollak et al.ICLR 2026 · 6 citations
