Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and Planning
Yizhe Huang, Anji Liu, Fanqi Kong, Yaodong Yang, Song-Chun Zhu, Xue Feng
Abstract
Despite the recent successes of multi-agent reinforcement learning (MARL) algorithms, efficiently adapting to co-players in mixed-motive environments remains a significant challenge. One feasible approach is to hierarchically model co-players' behavior based on inferring their characteristics. However, these methods often encounter difficulties in efficient reasoning and utilization of inferred information. To address these issues, we propose Hierarchical Opponent modeling and Planning (HOP), a novel multi-agent decision-making algorithm that enables few-shot adaptation to unseen policies in mixed-motive environments. HOP is hierarchically composed of two modules: an opponent modeling module that infers others' goals and learns corresponding goal-conditioned policies, and a planning module that employs Monte Carlo Tree Search (MCTS) to identify the best response. Our approach improves efficiency by updating beliefs about others' goals both across and within episodes and by using information from the opponent modeling module to guide planning. Experimental results demonstrate that in mixed-motive environments, HOP exhibits superior few-shot adaptation capabilities when interacting with various unseen agents, and excels in self-play scenarios. Furthermore, the emergence of social intelligence during our experiments underscores the potential of our approach in complex multi-agent environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3994dba-ce54-470b-9de7-024552b06566Cited by top-tier papers3
- Learning to Balance Altruism and Self-interest Based on Empathy in Mixed-Motive GamesFanqi Kong, Yizhe Huang, Song-Chun Zhu, Siyuan Qi et al.NeurIPS 2024 · 11 citations
- Social World Model-Augmented Mechanism Design Policy LearningXiaoyuan Zhang, Yizhe Huang, Chengdong Ma, Zhixun Chen et al.NeurIPS 2025 · 3 citations
- Planning with Quantized Opponent ModelsXiaopeng Yu, Kefan Su, Zongqing LuNeurIPS 2025
Builds on8
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand et al.ICML 2023 · 155 citations
- Model-Free Opponent ShapingChristopher Lu, Timon Willi, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 53 citations
- Risk-Averse Bayes-Adaptive Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2021 · 50 citations
- Dynamic Inverse Reinforcement Learning for Characterizing Animal BehaviorZoe Ashwood, Aditi Jha, Jonathan W. PillowNeurIPS 2022 · 50 citations
Related papers
- Opponent Modeling based on Subgoal InferenceXiaopeng Yu, Jiechuan Jiang, Zongqing LuNeurIPS 2024 · 7 citations
- Meta Reinforcement Learning with Autonomous Inference of Subtask DependenciesSungryull Sohn, Hyunjae Woo, Jongwook Choi, Honglak LeeICLR 2020 · 36 citations
- Model-Based Opponent ModelingXiaopeng Yu, Jiechuan Jiang, Wanpeng Zhang, Haobin Jiang et al.NeurIPS 2022 · 56 citations
- Efficient Meta Reinforcement Learning for Preference-based Fast AdaptationZhizhou Ren, Anji Liu, Yitao Liang, Jian Peng et al.NeurIPS 2022 · 11 citations
- Multi-Agent Actor-Critic with Hierarchical Graph Attention NetworkHeechang Ryu, Hayong Shin, Jinkyoo ParkAAAI 2020 · 143 citations
