Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach
Weiyu Ma, Qirui Mi, Yongcheng Zeng, Xue Yan, Runji Lin, Yuqiao Wu, Jun Wang, Haifeng Zhang
摘要
StarCraft II is a challenging benchmark for AI agents due to the necessity of both precise micro level operations and strategic macro awareness. Previous works, such as Alphastar and SCC, achieve impressive performance on tackling StarCraft II , however, still exhibit deficiencies in long term strategic planning and strategy interpretability. Emerging large language model (LLM) agents, such as Voyage and MetaGPT, presents the immense potential in solving intricate tasks. Motivated by this, we aim to validate the capabilities of LLMs on StarCraft II, a highly complex RTS game.To conveniently take full advantage of LLMs reasoning abilities, we first develop textual StratCraft II environment, called TextStarCraft II, which LLM agent can interact. Secondly, we propose a Chain of Summarization method, including single frame summarization for processing raw observations and multi frame summarization for analyzing game information, providing command recommendations, and generating strategic decisions. Our experiment consists of two parts: first, an evaluation by human experts, which includes assessing the LLMss mastery of StarCraft II knowledge and the performance of LLM agents in the game; second, the in game performance of LLM agents, encompassing aspects like win rate and the impact of Chain of Summarization.Experiment results demonstrate that: 1. LLMs possess the relevant knowledge and complex planning abilities needed to address StarCraft II scenarios; 2. Human experts consider the performance of LLM agents to be close to that of an average player who has played StarCraft II for eight years; 3. LLM agents are capable of defeating the built in AI at the Harder(Lv5) difficulty level. We have open sourced the code and released demo videos of LLM agent playing StarCraft II.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision MakingYangyang Yu, Zhiyuan Yao, Haohang Li, Zhiyang Deng 等NeurIPS 2024 · 被引用 197 次
- Language Agents with Reinforcement Learning for Strategic Play in the Werewolf GameZelai Xu, Chao Yu, Fei Fang, Yu Wang 等ICML 2024 · 被引用 145 次
- Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video GamesDongmin Park, Minkyu Kim, Beongjun Choi, Junhyuck Kim 等ICLR 2026 · 被引用 30 次
- Learning to Discuss Strategically: A Case Study on One Night Ultimate WerewolfXuanfa Jin, Ziyan Wang, Yali Du, Meng Fang 等NeurIPS 2024 · 被引用 28 次
- Mechanism Design for LLM Fine-tuning with Multiple Reward ModelsHaoran Sun, Yurong Chen, Siwei Wang, Chu Xu 等NeurIPS 2025 · 被引用 26 次
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk 等ICLR 2021 · 被引用 819 次
- Token-level Direct Preference OptimizationYongcheng Zeng, Guoqing Liu, Weiyu Ma, Ning Yang 等ICML 2024 · 被引用 136 次
相关 Paper
- LLM-PySC2: Starcraft II learning environment for Large Language ModelsZongyuan Li, Yanan Ni, Runnan Qi, Chang Lu 等NeurIPS 2025 · 被引用 15 次
- Large Language Models Are Neurosymbolic ReasonersMeng Fang, Shilong Deng, Yudi Zhang, Zijing Shi 等AAAI 2024 · 被引用 53 次
- Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of GoYichuan Ma, Linyang Li, Yongkang Chen, Peiji Li 等NeurIPS 2025 · 被引用 2 次
- GAMEBoT: Transparent Assessment of LLM Reasoning in GamesWenye Lin, Jonathan Roberts, Yunhan Yang, Samuel Albanie 等ACL 2025 · 被引用 12 次
- Agentic Video Summarization via Self-Reflecting Multimodal UnderstandingMiaotian Guo, Shuguang Dou, Yin Li, Aidong Men 等CVPR 2026
