Environment Generation for Zero-Shot Compositional Reinforcement Learning
Izzeddin Gur, Natasha Jaques, Yingjie Miao, Jongwook Choi, Manoj Tiwari, Honglak Lee, Aleksandra Faust
Abstract
Many real-world problems are compositional - solving them requires completing interdependent sub-tasks, either in series or in parallel, that can be represented as a dependency graph. Deep reinforcement learning (RL) agents often struggle to learn such complex tasks due to the long time horizons and sparse rewards. To address this problem, we present Compositional Design of Environments (CoDE), which trains a Generator agent to automatically build a series of compositional tasks tailored to the RL agent's current skill level. This automatic curriculum not only enables the agent to learn more complex tasks than it could have otherwise, but also selects tasks where the agent's performance is weak, enhancing its robustness and ability to generalize zero-shot to unseen tasks at test-time. We analyze why current environment generation techniques are insufficient for the problem of generating compositional tasks, and propose a new algorithm that addresses these issues. Our results assess learning and generalization across multiple compositional tasks, including the real-world problem of learning to navigate and interact with web pages. We learn to generate environments composed of multiple pages or rooms, and train RL agents capable of completing wide-range of complex tasks in those environments. We contribute two new benchmark frameworks for generating compositional tasks, compositional MiniGrid and gMiniWoB for web navigation.CoDE yields 4x higher success rate than the strongest baseline, and demonstrates strong performance of real websites learned on 3500 primitive tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 565333cd-ea44-44fa-aed4-0cab155e6b1eCited by top-tier papers10
- Language Models can Solve Computer TasksGeunwoo Kim, Pierre Baldi, Stephen McAleerNeurIPS 2023 · 539 citations
- Multimodal Web Navigation with Instruction-Finetuned Foundation ModelsHiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo et al.ICLR 2024 · 160 citations
- WebLINX: Real-World Website Navigation with Multi-Turn DialogueXing Han Lù, Zdenek Kasner, Siva ReddyICML 2024 · 146 citations
- Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer ControlLongtao Zheng, Rundong Wang, Xinrun Wang, Bo AnICLR 2024 · 132 citations
- A data-driven approach for learning to control computersPeter Conway Humphreys, David Raposo, Tobias Pohlen, Gregory Thornton et al.ICML 2022 · 124 citations
Builds on7
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- Enhanced POET: Open-ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their SolutionsRui Wang, Joel Lehman, Aditya Rawal, Jiale Zhi et al.ICML 2020 · 148 citations
- Automatic Curriculum Learning through Value DisagreementYunzhi Zhang, Pieter Abbeel, Lerrel PintoNeurIPS 2020 · 132 citations
- Learning with AMIGo: Adversarially Motivated Intrinsic GoalsAndres Campero, Roberta Raileanu, Heinrich Küttler, Joshua B. Tenenbaum et al.ICLR 2021 · 48 citations
- Meta Reinforcement Learning with Autonomous Inference of Subtask DependenciesSungryull Sohn, Hyunjae Woo, Jongwook Choi, Honglak LeeICLR 2020 · 36 citations
Related papers
- Beyond Fixed Tasks: Unsupervised Environment Design for Task-Level PairsDaniel Furelos-Blanco, Charles Pert, Frederik Kelbel, Alex F. Spies et al.AAAI 2026
- Synthetic Curriculum Reinforces Compositional Text-to-Image GenerationShijian Wang, Runhao Fu, Siyi Zhao, Qingqin Zhan et al.CVPR 2026 · 1 citation
- InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent TrainingZiyun Zhang, Zezhou Wang, Xiaoyi Zhang, Zongyu Guo et al.ACL 2026 · 11 citations
- ExeDec: Execution Decomposition for Compositional Generalization in Neural Program SynthesisKensen Shi, Joey Hong, Yinlin Deng, Pengcheng Yin et al.ICLR 2024 · 21 citations
- TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RLClément Romac, Rémy Portelas, Katja Hofmann, Pierre-Yves OudeyerICML 2021 · 28 citations
