Dreaming in Code for Curriculum Learning in Open-Ended Worlds
Konstantinos Mitsides, Maxence Faldor, Antoine Cully
摘要
Open-ended learning frames intelligence as emerging from continual interaction with an everexpanding space of environments. While recent advances have utilized foundation models to programmatically generate diverse environments, these approaches often focus on discovering isolated behaviors rather than orchestrating sustained progression. In complex open-ended worlds, the large combinatorial space of possible challenges makes it difficult for agents to discover sequences of experiences that remain consistently learnable. To address this, we propose Dreaming in Code (DiCode), a framework in which foundation models synthesize executable environment code to scaffold learning toward increasing competence. In DiCode, "dreaming" takes the form of materializing code-level variations of the world. We instantiate DiCode in Craftax, a challenging open-ended benchmark characterized by rich mechanics and long-horizon progression. Empirically, DiCode enables agents to acquire long-horizon skills, achieving a 16% improvement in mean return over the strongest baseline and non-zero success on late-game combat tasks where prior methods fail. Our results suggest that code-level environment design provides a practical mechanism for curriculum control, enabling the construction of intermediate environments that bridge competence gaps in open-ended worlds. Project page and source code are available at https://konstantinosmitsid es.github.io/dreaming-in-code and https://github.com/konstantinosm itsides/dreaming-in-code .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang 等ICLR 2024 · 被引用 582 次
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu 等ICML 2020 · 被引用 464 次
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu 等NeurIPS 2020 · 被引用 251 次
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 被引用 211 次
相关 Paper
- Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement LearningMichael T. Matthews, Michael Beukman, Benjamin Ellis, Mikayel Samvelyan 等ICML 2024 · 被引用 71 次
- Environment Generation for Zero-Shot Compositional Reinforcement LearningIzzeddin Gur, Natasha Jaques, Yingjie Miao, Jongwook Choi 等NeurIPS 2021 · 被引用 50 次
- CraftFactory: A Conditioned Control Policy Benchmark for Compositional GeneralizationJinbing Hou, Youpeng Zhao, Jian ZhaoAAAI 2025
- OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in CodeMaxence Faldor, Jenny Zhang, Antoine Cully, Jeff CluneICLR 2025 · 被引用 3 次
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding AgentsJingzhe Ding, Shengda Long, Changxin Pu, Ge Zhang 等ICML 2026 · 被引用 37 次
