Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning
Michael T. Matthews, Michael Beukman, Benjamin Ellis, Mikayel Samvelyan, Matthew Thomas Jackson, Samuel Coward, Jakob Nicolaus Foerster
摘要
Benchmarks play a crucial role in the development and analysis of reinforcement learning (RL) algorithms. We identify that existing benchmarks used for research into open-ended learning fall into one of two categories. Either they are too slow for meaningful research to be performed without enormous computational resources, like Crafter, NetHack and Minecraft, or they are not complex enough to pose a significant challenge, like Minigrid and Procgen. To remedy this, we first present Craftax-Classic: a ground-up rewrite of Crafter in JAX that runs up to 250x faster than the Python-native original. A run of PPO using 1 billion environment interactions finishes in under an hour using only a single GPU and averages 90% of the optimal reward. To provide a more compelling challenge we present the main Craftax benchmark, a significant extension of the Crafter mechanics with elements inspired from NetHack 1 . Solving Craftax requires deep exploration, long term planning and memory, as well as continual adaptation to novel situations as more of the world is discovered. We show that existing methods including global and episodic exploration, as well as unsupervised environment design fail to make material progress on the benchmark. We believe that Craftax can for the first time allow researchers to experiment in a complex, open-ended environment with limited computational resources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAXClément Bonnet, Daniel Luo, Donal Byrne, Shikha Surana 等ICLR 2024 · 被引用 52 次
- No Regrets: Investigating and Improving Regret Approximations for Curriculum DiscoveryAlexander Rutherford, Michael Beukman, Timon Willi, Bruno Lacerda 等NeurIPS 2024 · 被引用 38 次
- Can Learned Optimization Make Reinforcement Learning Less Difficult?Alexander David Goldie, Chris Lu, Matthew Thomas Jackson, Shimon Whiteson 等NeurIPS 2024 · 被引用 18 次
- Evolution Strategies at the HyperscaleBidipta Sarkar, Mattie Fellows, Juan Duque, Alistair Letcher 等ICML 2026 · 被引用 16 次
- Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam TimestepsBenjamin Ellis, Matthew Thomas Jackson, Andrei Lupu, Alexander David Goldie 等NeurIPS 2024 · 被引用 14 次
它引用的顶会 Paper12
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu 等NeurIPS 2020 · 被引用 251 次
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 被引用 211 次
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides 等ICLR 2020 · 被引用 204 次
相关 Paper
- Benchmarking the Spectrum of Agent CapabilitiesDanijar HafnerICLR 2022 · 被引用 193 次
- Unsupervised Hierarchical Skill DiscoveryDamion Harvey, Geraud Nangue Tasse, Benjamin Rosman, Branden Ingram 等ICML 2026 · 被引用 1 次
- Dreaming in Code for Curriculum Learning in Open-Ended WorldsKonstantinos Mitsides, Maxence Faldor, Antoine CullyICML 2026 · 被引用 1 次
- Craftium: Bridging Flexibility and Efficiency for Rich 3D Single- and Multi-Agent EnvironmentsMikel Malagón, Josu Ceberio, José Antonio LozanoICML 2025
- Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAXWaris Radji, Thomas Michel, Hector PiteauICLR 2026 · 被引用 5 次
