Unsupervised Training Sequence Design: Efficient and Generalizable Agent Training
Wenjun Li, Pradeep Varakantham
Abstract
To train generalizable Reinforcement Learning (RL) agents, researchers recently proposed the Unsupervised Environment Design (UED) framework, in which a teacher agent creates a very large number of training environments and a student agent trains on the experiences in these environments to be robust against unseen testing scenarios. For example, to train a student to master the "stepping over stumps" task, the teacher will create numerous training environments with varying stump heights and shapes. In this paper, we argue that UED neglects training efficiency and its need for a very large number of environments (henceforth referred to as infinite horizon training) makes it less suitable for training robots and non-expert humans. In real-world applications where either creating new training scenarios is expensive or training efficiency is of critical importance, we want to maximize both the learning efficiency and learning outcome of the student. To achieve efficient finite horizon training, we propose a novel Markov Decision Process (MDP) formulation for the teacher agent, referred to as Unsupervised Training Sequence Design (UTSD). Specifically, we encode salient information from the student policy (e.g., behaviors and learning progress) into the teacher's state space, enabling the teacher to closely track the student's learning progress and consequently discover the optimal training sequences with finite lengths. Additionally, we explore the teacher's efficient adaptation to unseen students at test time by employing the context-based metalearning approach, which leverages the teacher's past experiences with various students. Finally, we empirically demonstrate our teacher's capability to design efficient and effective training sequences for students with varying capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81d8bf0c-f5ea-49f9-abf6-5dd3bea97727Builds on9
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 211 citations
- Evolving Curricula with Regret-Based Environment DesignJack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan et al.ICML 2022 · 175 citations
Related papers
- Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed GoalsOctavio PappalardoICLR 2026
- Marginal Benefit Driven RL Teacher for Unsupervised Environment DesignDexun Li, Wenjun Li, Pradeep VarakanthamAAAI 2025 · 1 citation
- Improving Environment Novelty Quantification for Effective Unsupervised Environment DesignJayden Teoh, Wenjun Li, Pradeep VarakanthamNeurIPS 2024 · 6 citations
- TRACED: Transition-aware Regret Approximation with Co-learnability for Environment DesignGeonwoo Cho, Jaegyun Im, Jihwan Lee, Hojun Yi et al.ICLR 2026 · 1 citation
- Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementZhi Wang, Li Zhang, Wenhao Wu, Yuanheng Zhu et al.NeurIPS 2024 · 31 citations
