Fast Peer Adaptation with Context-aware Exploration
Long Ma, Yuanfei Wang, Fangwei Zhong, Song-Chun Zhu, Yizhou Wang
Abstract
Fast adapting to unknown peers (partners or opponents) with different strategies is a key challenge in multi-agent games. To do so, it is crucial for the agent to probe and identify the peer's strategy efficiently, as this is the prerequisite for carrying out the best response in adaptation. However, exploring the strategies of unknown peers is difficult, especially when the games are partially observable and have a long horizon. In this paper, we propose a peer identification reward, which rewards the learning agent based on how well it can identify the behavior pattern of the peer over the historical context, such as the observation over multiple episodes. This reward motivates the agent to learn a context-aware policy for effective exploration and fast adaptation, i.e., to actively seek and collect informative feedback from peers when uncertain about their policies and to exploit the context to perform the best response when confident. We evaluate our method on diverse testbeds that involve competitive (Kuhn Poker), cooperative (PO-Overcooked), or mixed (Predator-Prey-W) games with peer agents. We demonstrate that our method induces more active exploration behavior, achieving faster adaptation and better outcomes than existing methods 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed8bbe0e-5cce-4aad-8d73-3988437736aeCited by top-tier papers9
- Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyZhenyu Guan, Xiangyu Kong, Fangwei Zhong, Yizhou WangNeurIPS 2024 · 48 citations
- Fine Tuning Out-of-Vocabulary Item Recommendation with User Sequence ImaginationRuochen Liu, Hao Chen, Yuanchen Bei, Qijie Shen et al.NeurIPS 2024 · 22 citations
- Adaptively Coordinating with Novel Partners via Learned Latent StrategiesBenjamin Li, Shuyang Shi, Lucia Romero, Huao Li et al.NeurIPS 2025 · 4 citations
- Offline Opponent Modeling with Truncated Q-driven Instant Policy RefinementYuheng Jing, Kai Li, Bingyun Liu, Ziwen Zhang et al.ICML 2025
- CooT: Learning to Coordinate In-Context with Coordination TransformersHuai-Chih Wang, Hsiang-Chun Chuang, Hsi-Chun Cheng, Dai-Jie Wu et al.ICML 2026
Builds on20
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo et al.ICLR 2020 · 349 citations
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes et al.NeurIPS 2021 · 239 citations
- Trajectory Diversity for Zero-Shot CoordinationAndrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. FoersterICML 2021 · 157 citations
- Cooperative Exploration for Multi-Agent Deep Reinforcement LearningIou-Jen Liu, Unnat Jain, Raymond A. Yeh, Alexander G. SchwingICML 2021 · 133 citations
Related papers
- AIR: Unifying Individual and Collective Exploration in Cooperative Multi-Agent Reinforcement LearningGuangchong Zhou, Zeren Zhang, Guoliang FanAAAI 2025
- Peer Learning: Learning Complex Policies in Groups from Scratch via Action RecommendationsCedric Derstroff, Mattia Cerrato, Jannis Brugger, Jan Peters et al.AAAI 2024 · 1 citation
- Learning to Collaborate with Unknown Agents in the Absence of RewardZuyuan Zhang, Hanhan Zhou, Mahdi Imani, Taeyoung Lee et al.AAAI 2025 · 11 citations
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen et al.ICML 2021 · 33 citations
- Opponent Modeling based on Subgoal InferenceXiaopeng Yu, Jiechuan Jiang, Zongqing LuNeurIPS 2024 · 7 citations
