Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
Dongmin Park, Minkyu Kim, Beongjun Choi, Junhyuck Kim, Keon Lee, Jonghyun Lee, Inkyu Park, Byeong-Uk Lee, Jaeyoung Hwang, Jaewoo Ahn, Ameya Mahabaleshwarkar, Bilal Kartal
Abstract
Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game benchmarks fall short of practical needs: they lack evaluations of diverse LLM capabilities across various game genres, studies of agentic modules crucial for complex gameplay, and fine-tuning datasets to adapt pre-trained LLMs into gaming agents. To fill these gaps, we present Orak, a benchmark for training and evaluating LLM agents across 12 popular video games spanning all major genres. Using a plug-and-play interface built on Model Context Protocol (MCP), Orak supports systematic and reproducible studies of agentic modules in varied game scenarios. We further release a fine-tuning dataset of expert LLM gameplay trajectories spanning multiple genres, turning general LLMs into effective game agents. Orak offers a comprehensive evaluation framework, including game leaderboards, LLM battle arenas, and in-depth analyses of input modality, agentic strategies, and fine-tuning effects, establishing a foundation towards versatile gaming agents. Code and datasets are available at https://github.com/krafton-ai/Orak and https://huggingface.co/datasets/KRAFTON/Orak.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ce5dea3-23b4-439a-81a8-48aa8144bda0Cited by top-tier papers6
- Meta-RL Induces Exploration in Language AgentsYulun Jiang, Liangze Jiang, Damien Teney, Michael Moor et al.ICLR 2026 · 20 citations
- AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and ProgressZhiheng Xi, Chenyang Liao, Guanyu Li, Zhihao Zhang et al.WWW 2026 · 19 citations
- VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?Minkyu Kim, Sangheon Lee, Dongmin ParkICLR 2026 · 5 citations
- AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World ScenariosLisa Alazraki, Lihu Chen, Ana Brassard, Joe Stacey et al.ACL 2026 · 1 citation
- Investigating Advanced Reasoning of Large Language Models via Black-Box Environment InteractionCongchi Yin, Tianyi Wu, Yankai Shu, Alex Gu et al.ICML 2026 · 1 citation
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
Related papers
- SmartPlay : A Benchmark for LLMs as Intelligent AgentsYue Wu, Xuan Tang, Tom M. Mitchell, Yuanzhi LiICLR 2024 · 121 citations
- clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational AgentsKranti Chalamalasetti, Jana Götze, Sherzod Hakimov, Brielen Madureira et al.EMNLP 2023 · 6 citations
- TritonGym: A Benchmark for Agentic LLM Workflows in Triton GPU Code GenerationYue Guan, Yichen Lin, Xu Zhao, Jianzhu Yao et al.ICML 2026
- AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse EnvironmentsZhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong et al.ACL 2025 · 20 citations
- BALROG: Benchmarking Agentic LLM and VLM Reasoning On GamesDavide Paglieri, Bartlomiej Cupial, Samuel Coward, Ulyana Piterbarg et al.ICLR 2025
