AIR: Unifying Individual and Collective Exploration in Cooperative Multi-Agent Reinforcement Learning
Guangchong Zhou, Zeren Zhang, Guoliang Fan
摘要
Exploration in cooperative multi-agent reinforcement learning (MARL) remains challenging for value-based agents due to the absence of an explicit policy. Existing approaches include individual exploration based on uncertainty towards the system and collective exploration through behavioral diversity among agents. However, the introduction of additional structures often leads to reduced training efficiency and infeasible integration of these methods. In this paper, we propose Adaptive exploration via Identity Recognition (AIR), which consists of two adversarial components: a classifier that recognizes agent identities from their trajectories, and an action selector that adaptively adjusts the mode and degree of exploration. We theoretically prove that AIR can facilitate both individual and collective exploration during training, and experiments also demonstrate the efficiency and effectiveness of AIR across various tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac 等AAAI 2020 · 被引用 496 次
- Celebrating Diversity in Shared Multi-Agent Reinforcement LearningChenghao Li, Tonghan Wang, Chengjie Wu, Qianchuan Zhao 等NeurIPS 2021 · 被引用 224 次
- Influence-Based Multi-Agent ExplorationTonghan Wang, Jianhao Wang, Yi Wu, Chongjie ZhangICLR 2020 · 被引用 156 次
- Episodic Multi-agent Reinforcement Learning with Curiosity-driven ExplorationLulu Zheng, Jiarui Chen, Jianhao Wang, Jiamin He 等NeurIPS 2021 · 被引用 126 次
相关 Paper
- HyperMARL: Adaptive Hypernetworks for Multi-Agent RLKale-ab Abebe Tessera, Arrasy Rahman, Amos J. Storkey, Stefano V. AlbrechtNeurIPS 2025 · 被引用 11 次
- HCPO: Hierarchical Conductor-Based Policy Optimization in Multi-Agent Reinforcement LearningZejiao Liu, Junqi Tu, Yitian Hong, Luolin Xiong 等AAAI 2026
- UneVEn: Universal Value Exploration for Multi-Agent Reinforcement LearningTarun Gupta, Anuj Mahajan, Bei Peng, Wendelin Boehmer 等ICML 2021 · 被引用 59 次
- Two Heads are Better Than One: A Simple Exploration Framework for Efficient Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Xingchen Li 等NeurIPS 2023 · 被引用 7 次
- Toward Efficient Multi-Agent Exploration With Trajectory Entropy MaximizationTianxu Li, Kun ZhuICLR 2025
