Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent Intelligence
Yuhang Song, Andrzej Wojcicki, Thomas Lukasiewicz, Jianyi Wang, Abi Aryan, Zhenghua Xu, Mai Xu, Zihan Ding, Lianlong Wu
摘要
Learning agents that are not only capable of taking tests, but also innovating is becoming a hot topic in AI. One of the most promising paths towards this vision is multi-agent learning, where agents act as the environment for each other, and improving each agent means proposing new problems for others. However, existing evaluation platforms are either not compatible with multi-agent settings, or limited to a specific game. That is, there is not yet a general evaluation platform for research on multi-agent intelligence. To this end, we introduce Arena, a general evaluation platform for multi-agent intelligence with 35 games of diverse logics and representations. Furthermore, multi-agent intelligence is still at the stage where many problems remain unexplored. Therefore, we provide a building toolkit for researchers to easily invent and build novel multi-agent problems from the provided game set based on a GUI-configurable social tree and five basic multi-agent reward schemes. Finally, we provide Python implementations of five state-of-the-art deep multi-agent reinforcement learning baselines. Along with the baseline implementations, we release a set of 100 best agents/teams that we can train with different training schemes for each game, as the base for evaluating agents with population performance. As such, the research community can perform comparisons under a stable and uniform standard. All the implementations and accompanied tutorials have been open-sourced for the community at https://sites.google.com/view/arena-unity/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting PotJoel Z. Leibo, Edgar A. Duéñez-Guzmán, Alexander Vezhnevets, John P. Agapiou 等ICML 2021 · 被引用 134 次
- Learning to Play No-Press Diplomacy with Best Response Policy IterationThomas W. Anthony, Tom Eccles, Andrea Tacchetti, János Kramár 等NeurIPS 2020 · 被引用 50 次
- Dealing with Non-Stationarity in MARL via Trust-Region DecompositionWenhao Li, Xiangfeng Wang, Bo Jin, Junjie Sheng 等ICLR 2022 · 被引用 14 次
- FightLadder: A Benchmark for Competitive Multi-Agent Reinforcement LearningWenzhe Li, Zihan Ding, Seth Karten, Chi JinICML 2024 · 被引用 13 次
- SAPIEN: A SimulAted Part-Based Interactive ENvironmentFanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia 等CVPR 2020
相关 Paper
- Benchmarking Agent Memory in Interdependent Multi-Session Agentic TasksZexue He, Yu Wang, Churan Zhi, Yuanzhe Hu 等ICML 2026 · 被引用 2 次
- GoBigger: A Scalable Platform for Cooperative-Competitive Multi-Agent Interactive SimulationMing Zhang, Shenghan Zhang, Zhenjie Yang, Lekai Chen 等ICLR 2023
- DR-Arena: an Automated Evaluation Framework for Deep Research AgentsYiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan ZhangACL 2026 · 被引用 2 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Computer Agent Arena: Toward Human-Centric Evaluation and Analysis of Computer-Use AgentsBowen Wang, Xinyuan Wang, Jiaqi Deng, Tianbao Xie 等ICLR 2026
