FightLadder: A Benchmark for Competitive Multi-Agent Reinforcement Learning
Wenzhe Li, Zihan Ding, Seth Karten, Chi Jin
Abstract
Recent advances in reinforcement learning (RL) heavily rely on a variety of well-designed benchmarks, which provide environmental platforms and consistent criteria to evaluate existing and novel algorithms. Specifically, in multi-agent RL (MARL), a plethora of benchmarks based on cooperative games have spurred the development of algorithms that improve the scalability of cooperative multi-agent systems. However, for the competitive setting, a lightweight and open-sourced benchmark with challenging gaming dynamics and visual inputs has not yet been established. In this work, we present FightLadder, a real-time fighting game platform, to empower competitive MARL research. Along with the platform, we provide implementations of state-of-the-art MARL algorithms for competitive games, as well as a set of evaluation metrics to characterize the performance and exploitability of agents. We demonstrate the feasibility of this platform by training a general agent that consistently defeats 12 built-in characters in single-player mode, and expose the difficulty of training a non-exploitable agent without human knowledge and demonstrations in two-player mode. FightLadder provides meticulously designed environments to address critical challenges in competitive MARL research, aiming to catalyze a new era of discovery and advancement in the field. Videos and code at https://sites.google.com/view/fightladder/home.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a38be593-fee0-42c9-809c-0e235f919a6eCited by top-tier papers2
- Sparta Alignment: Collectively Aligning Multiple Language Models through CombatYuru Jiang, Wenxuan Ding, Shangbin Feng, Greg Durrett et al.NeurIPS 2025 · 8 citations
- PokéChamp: an Expert-level Minimax Language AgentSeth Karten, Andy Luu Nguyen, Chi JinICML 2025
Builds on7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac et al.AAAI 2020 · 496 citations
- PettingZoo: Gym for Multi-Agent Reinforcement LearningJ. K. Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar et al.NeurIPS 2021 · 478 citations
- Independent Policy Gradient Methods for Competitive Reinforcement LearningConstantinos Daskalakis, Dylan J. Foster, Noah GolowichNeurIPS 2020 · 200 citations
Related papers
- CAMAR: Continuous Actions Multi-Agent RoutingArtem Pshenitsyn, Aleksandr Panov, Alexey SkrynnikAAAI 2026 · 2 citations
- GoBigger: A Scalable Platform for Cooperative-Competitive Multi-Agent Interactive SimulationMing Zhang, Shenghan Zhang, Zhenjie Yang, Lekai Chen et al.ICLR 2023
- Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent IntelligenceYuhang Song, Andrzej Wojcicki, Thomas Lukasiewicz, Jianyi Wang et al.AAAI 2020 · 36 citations
- TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement LearningHayeong Lee, JunHyeok Oh, Byung-Jun LeeICML 2026
- MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement LearningYi Wang, Ningze Zhong, Zhiheng Fu, Longguang Wang et al.CVPR 2026
