Agents in the Sandbox: End-to-End Crash Bug Reproduction for Minecraft
Eray Yapagci, Yavuz Alp Sencer Öztürk, Eray Tüzün
Abstract
Reproducing game bugs, particularly crash bugs in continuously evolving games like Minecraft, is a notoriously manual, time-consuming, and challenging process to automate; insights from a key decision maker from Minecraft we interviewed confirm this, highlighting that a substantial portion of crash reports necessitate manual scenario reconstruction. Despite the success of LLM-driven bug reproduction in other software domains, games, with their complex interactive environments, remain largely unaddressed. This paper introduces BugCraft, a novel end-to-end framework designed to automate the reproduction of crash bugs in Minecraft directly from user-submitted bug reports, addressing the critical gap in automated game bug reproduction. BugCraft employs a two-stage approach: first, a Step Synthesizer leverages LLMs and Minecraft Wiki knowledge to transform bug reports into high-quality, structured steps to reproduce (S2R). Second, an Action Model, powered by a vision-based LLM agent and a custom macro API, executes these S2R steps within Minecraft to trigger the reported crash. To facilitate evaluation, we introduce BugCraft-Bench, a curated dataset of Minecraft crash bug reports. On BugCraft-Bench, our framework end-to-end reproduced 34.9% of crash bugs with GPT-4.1, outperforming baseline computer-use models by 37%. BugCraft demonstrates the feasibility of automated reproduction of crash bugs in complex game environments using LLMs, opening promising avenues for game testing and development. Finally, we make our code open at https://bugcraft2025.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9562fc3a-9af0-4d48-b0d7-91178560f474Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Large Language Models are Few-shot Testers: Exploring LLM-based General Bug ReproductionSungmin Kang, Juyeon Yoon, Shin YooICSE 2023 · 163 citations
Related papers
- Automatically Reproducing Android Bug Reports using Natural Language Processing and Reinforcement LearningZhaoxu Zhang, Robert Winn, Yu Zhao, Tingting Yu et al.ISSTA 2023 · 14 citations
- SPRING: Studying Papers and Reasoning to play GamesYue Wu, So Yeon Min, Shrimai Prabhumoye, Yonatan Bisk et al.NeurIPS 2023 · 32 citations
- DBugScribe: Automatic Database Bug Reproduction from Community ReportsSuyang Zhong, Mo Sha, Sheng Wang, Fangyuan Zhou et al.SIGMOD 2026 · 1 citation
- Feedback-Driven Automated Whole Bug Report Reproduction for Android AppsDingbang Wang, Yu Zhao, Sidong Feng, Zhaoxu Zhang et al.ISSTA 2024 · 16 citations
- Synthetic Repo-level Bug Dataset for Training Automated Program Repair ModelsMinh V. T. Pham, Huy N. Phan, Nhat Hoang Phan, Cuong Chi Le et al.ICSE 2026
