TRACED: Transition-aware Regret Approximation with Co-learnability for Environment Design
Geonwoo Cho, Jaegyun Im, Jihwan Lee, Hojun Yi, Sejin Kim, Sundong Kim
摘要
Generalizing deep reinforcement learning agents to unseen environments remains a significant challenge. One promising solution is Unsupervised Environment Design (UED), a co-evolutionary framework in which a teacher adaptively generates tasks with high learning potential, while a student learns a robust policy from this evolving curriculum. Existing UED methods typically measure learning potential via regret, the gap between optimal and current performance, approximated solely by value-function loss. Building on these approaches, we introduce the transition-prediction error as an additional term in our regret approximation. To capture how training on one task affects performance on others, we further propose a lightweight metric called Co-Learnability. By combining these two measures, we present Transition-aware Regret Approximation with Co-learnability for Environment Design (TRACED). Empirical evaluations show that TRACED produces curricula that improve zero-shot generalization over strong baselines across multiple benchmarks. Ablation studies confirm that the transition-prediction error drives rapid complexity ramp-up and that Co-Learnability delivers additional gains when paired with the transition-prediction error. These results demonstrate how refined regret approximation and explicit modeling of task relationships can be leveraged for sample-efficient curriculum design in UED. Project Page: https://geonwoo.me/traced/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 被引用 211 次
- Evolving Curricula with Regret-Based Environment DesignJack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan 等ICML 2022 · 被引用 175 次
- Enhanced POET: Open-ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their SolutionsRui Wang, Joel Lehman, Aditya Rawal, Jiale Zhi 等ICML 2020 · 被引用 148 次
相关 Paper
- Improving Regret Approximation for Unsupervised Dynamic Environment GenerationHarry Mead, Bruno Lacerda, Jakob N. Foerster, Nick HawesNeurIPS 2025 · 被引用 1 次
- Adversarial Environment Design via Regret-Guided Diffusion ModelsHojun Chung, Junseo Lee, Minsoo Kim, Dohyeong Kim 等NeurIPS 2024 · 被引用 11 次
- PACE: Parameter Change for Unsupervised Environment DesignFang YUAN, Junjie Zeng, Qinglun Li, Long Qin 等ICML 2026
- CLUTR: Curriculum Learning via Unsupervised Task Representation LearningAbdus Salam Azad, Izzeddin Gur, Jasper Emhoff, Nathaniel Alexis 等ICML 2023 · 被引用 20 次
- Improving Environment Novelty Quantification for Effective Unsupervised Environment DesignJayden Teoh, Wenjun Li, Pradeep VarakanthamNeurIPS 2024 · 被引用 6 次
