SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social Dilemmas
Zihao Guo, Shuqing Shi, Richard Willis, Tristan Tomilin, Joel Z. Leibo, Yali Du
Abstract
Sequential social dilemmas pose a significant challenge in the field of multi-agent reinforcement learning (MARL), requiring environments that accurately reflect the tension between individual and collective interests. Previous benchmarks and environments, such as Melting Pot, provide an evaluation protocol that measures generalization to new social partners in various test scenarios. However, running reinforcement learning algorithms in traditional environments requires substantial computational resources. In this paper, we introduce SocialJax, a suite of sequential social dilemma environments and algorithms implemented in JAX. JAX is a high-performance numerical computing library for Python that enables significant improvements in operational efficiency. Our experiments demonstrate that the SocialJax training pipeline achieves at least 50× speed-up in real-time performance compared to Melting Pot's RLlib baselines. Additionally, we validate the effectiveness of baseline algorithms within SocialJax environments. Finally, we use Schelling diagrams to verify the social dilemma properties of these environments, ensuring that they accurately capture the dynamics of social dilemmas. Our code is available at https://github.com/cooperativex/SocialJax .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0bfe0e7-af8f-42a4-8157-d91e7ac08469Cited by top-tier papers2
- MEAL: A Benchmark for Continual Multi-Agent Reinforcement LearningTristan Tomilin, Luka van den Boogaard, Samuel Garcin, Constantin Ruhdorfer et al.ICML 2026 · 9 citations
- Fair Cooperation in Mixed-Motive Games via Conflict-Aware Gradient AdjustmentWoojun Kim, Katia SycaraNeurIPS 2025 · 4 citations
Builds on8
- PettingZoo: Gym for Multi-Agent Reinforcement LearningJ. K. Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar et al.NeurIPS 2021 · 478 citations
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides et al.ICLR 2020 · 204 citations
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous ControlH. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark et al.ICLR 2020 · 138 citations
- Discovered Policy OptimisationChris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz et al.NeurIPS 2022 · 134 citations
- Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting PotJoel Z. Leibo, Edgar A. Duéñez-Guzmán, Alexander Vezhnevets, John P. Agapiou et al.ICML 2021 · 134 citations
Related papers
- TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement LearningHayeong Lee, JunHyeok Oh, Byung-Jun LeeICML 2026
- Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAXClément Bonnet, Daniel Luo, Donal Byrne, Shikha Surana et al.ICLR 2024 · 52 citations
- Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAXWaris Radji, Thomas Michel, Hector PiteauICLR 2026 · 5 citations
- Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement LearningMichael T. Matthews, Michael Beukman, Benjamin Ellis, Mikayel Samvelyan et al.ICML 2024 · 71 citations
- MAFE: Enabling Equitable Algorithm Design in Multi-Agent Multi-Stage Decision-Making SystemsZachary Lazri, Anirudh Nakra, Ivan Brugere, Danial Dervovic et al.ICML 2026
