Acquiring Diverse Skills using Curriculum Reinforcement Learning with Mixture of Experts
Onur Celik, Aleksandar Taranovic, Gerhard Neumann
摘要
Reinforcement learning (RL) is a powerful approach for acquiring a good-performing policy. However, learning diverse skills is challenging in RL due to the commonly used Gaussian policy parameterization. We propose Diverse Skill Learning (Di-SkilL), an RL method for learning diverse skills using Mixture of Experts, where each expert formalizes a skill as a contextual motion primitive. Di-SkilL optimizes each expert and its associate context distribution to a maximum entropy objective that incentivizes learning diverse skills in similar contexts. The per-expert context distribution enables automatic curricula learning, allowing each expert to focus on its best-performing sub-region of the context space. To overcome hard discontinuities and multi-modalities without any prior knowledge of the environment's unknown context probability space, we leverage energy-based models to represent the per-expert context distributions and demonstrate how we can efficiently train them using the standard policy gradient objective. We show on challenging robot simulation tasks that Di-SkilL can learn diverse and performant skills.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Diversity-Incentivized Exploration for Versatile ReasoningZican Hu, Shilin Zhang, Yafu Li, Jianhao Yan 等ICLR 2026 · 被引用 32 次
- Variational Distillation of Diffusion Policies into Mixture of ExpertsHongyi Zhou, Denis Blessing, Ge Li, Onur Celik 等NeurIPS 2024 · 被引用 19 次
- MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity RecognitionHao Zhang, Zhan Zhuang, Xuehao Wang, Xiaodong Yang 等NeurIPS 2025 · 被引用 11 次
- TROLL: Trust Regions Improve Reinforcement Learning for Large Language ModelsPhilipp Becker, Niklas Freymuth, Serge Thilges, Fabian Otto 等ICLR 2026 · 被引用 8 次
- Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMsYifan Zhou, Sachin Grover, Mohamed El Mistiri, Kamalesh Kalirathinam 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper19
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Behavior Transformers: Cloning modes with one stoneNur Muhammad Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, Lerrel PintoNeurIPS 2022 · 被引用 470 次
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 被引用 211 次
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering SkillsVictor Campos, Alexander Trott, Caiming Xiong, Richard Socher 等ICML 2020 · 被引用 178 次
相关 Paper
- Information Maximizing Curriculum: A Curriculum-Based Approach for Learning Versatile SkillsDenis Blessing, Onur Celik, Xiaogang Jia, Moritz Reuss 等NeurIPS 2023 · 被引用 19 次
- Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous GraspingZiye Huang, Haoqi Yuan, Yuhui Fu, Zongqing LuICLR 2025
- Learning a family of motor skills from a single motion clipSeyoung Lee, Sunmin Lee, Yongwoo Lee, Jehee LeeSIGGRAPH 2021 · 被引用 35 次
- Robust Policy Learning via Offline Skill DiffusionWoo Kyung Kim, Minjong Yoo, Honguk WooAAAI 2024 · 被引用 9 次
- Behavior Contrastive Learning for Unsupervised Skill DiscoveryRushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li 等ICML 2023 · 被引用 34 次
