Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
Haochen Sun, Shuwen Zhang, Lujie Niu, Lei Ren, Hao Xu, Hao Fu, Fangkun Zhao, Caixia Yuan, Xiaojie Wang
摘要
Large Language Models (LLMs) based agent systems have made great strides in real-world applications beyond traditional NLP tasks. This paper proposes a new LLM-based Multi-Agent System (LLM-MAS) benchmark, Collab-Overcooked, built on the popular Overcooked-AI game with more applicable and challenging tasks in interactive environments. Collab-Overcooked extends existing benchmarks in two novel ways. First, it provides a multi-agent framework supporting diverse tasks and objectives and encourages collaboration through natural language communication. Second, it introduces a spectrum of process-oriented evaluation metrics to assess the fine-grained collaboration capabilities of different LLM agents, a dimension often overlooked in prior work. We conduct extensive experiments with 13 popular LLMs and show that, while the LLMs exhibit a strong ability in goal interpretation, there are significant shortcomings in active collaboration and continuous adaptation, which are critical for efficiently fulfilling complex tasks. Notably, we highlight the strengths and weaknesses of LLM-MAS and provide insights for improving and evaluating LLM-MAS on a unified and open-source benchmark. The environments, 30 open-ended tasks, and the evaluation package are publicly available at https://github.com/YusaeMeow/Collab-Overcooked.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMsYuxuan Li, Aoi Naito, Hirokazu ShiradoICML 2026 · 被引用 9 次
- CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive EngagementHong Qian, Yuanhao Liu, Zihan Zhou, Zongbao Zhang 等ICML 2026
它引用的顶会 Paper7
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- ExpeL: LLM Agents Are Experiential LearnersAndrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin 等AAAI 2024 · 被引用 484 次
- Building Cooperative Embodied Agents Modularly with Large Language ModelsHongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou 等ICLR 2024 · 被引用 303 次
- MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue ResolutionWei Tao, Yucheng Zhou, Yanlin Wang, Wenqiang Zhang 等NeurIPS 2024 · 被引用 210 次
相关 Paper
- Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile AgentsShihan Deng, Weikai Xu, Hongda Sun, Wei Liu 等ACL 2024 · 被引用 10 次
- clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational AgentsKranti Chalamalasetti, Jana Götze, Sherzod Hakimov, Brielen Madureira 等EMNLP 2023 · 被引用 6 次
- ProAgent: Building Proactive Cooperative Agents with Large Language ModelsCeyao Zhang, Kaijie Yang, Siyi Hu, Zihao Wang 等AAAI 2024 · 被引用 141 次
- MultiAgentBench : Evaluating the Collaboration and Competition of LLM agentsKunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang 等ACL 2025 · 被引用 97 次
- LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language ModelsMarwa Abdulhai, Isadora White, Charlie Victor Snell, Charles Sun 等ICML 2025
