Unsupervised Training Sequence Design: Efficient and Generalizable Agent Training
Wenjun Li, Pradeep Varakantham
摘要
To train generalizable Reinforcement Learning (RL) agents, researchers recently proposed the Unsupervised Environment Design (UED) framework, in which a teacher agent creates a very large number of training environments and a student agent trains on the experiences in these environments to be robust against unseen testing scenarios. For example, to train a student to master the "stepping over stumps" task, the teacher will create numerous training environments with varying stump heights and shapes. In this paper, we argue that UED neglects training efficiency and its need for a very large number of environments (henceforth referred to as infinite horizon training) makes it less suitable for training robots and non-expert humans. In real-world applications where either creating new training scenarios is expensive or training efficiency is of critical importance, we want to maximize both the learning efficiency and learning outcome of the student. To achieve efficient finite horizon training, we propose a novel Markov Decision Process (MDP) formulation for the teacher agent, referred to as Unsupervised Training Sequence Design (UTSD). Specifically, we encode salient information from the student policy (e.g., behaviors and learning progress) into the teacher's state space, enabling the teacher to closely track the student's learning progress and consequently discover the optimal training sequences with finite lengths. Additionally, we explore the teacher's efficient adaptation to unseen students at test time by employing the context-based metalearning approach, which leverages the teacher's past experiences with various students. Finally, we empirically demonstrate our teacher's capability to design efficient and effective training sequences for students with varying capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 被引用 211 次
- Evolving Curricula with Regret-Based Environment DesignJack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan 等ICML 2022 · 被引用 175 次
相关 Paper
- Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed GoalsOctavio PappalardoICLR 2026
- Marginal Benefit Driven RL Teacher for Unsupervised Environment DesignDexun Li, Wenjun Li, Pradeep VarakanthamAAAI 2025 · 被引用 1 次
- Improving Environment Novelty Quantification for Effective Unsupervised Environment DesignJayden Teoh, Wenjun Li, Pradeep VarakanthamNeurIPS 2024 · 被引用 6 次
- TRACED: Transition-aware Regret Approximation with Co-learnability for Environment DesignGeonwoo Cho, Jaegyun Im, Jihwan Lee, Hojun Yi 等ICLR 2026 · 被引用 1 次
- Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementZhi Wang, Li Zhang, Wenhao Wu, Yuanheng Zhu 等NeurIPS 2024 · 被引用 31 次
