Meta-Reinforcement Learning via Exploratory Task Clustering
Zhendong Chu, Renqin Cai, Hongning Wang
Abstract
Meta-reinforcement learning (meta-RL) aims to quickly solve new RL tasks by leveraging knowledge from prior tasks. Previous studies often assume a single-mode homogeneous task distribution, ignoring possible structured heterogeneity among tasks. Such an oversight can hamper effective exploration and adaptation, especially with limited samples. In this work, we harness the structured heterogeneity among tasks via clustering to improve meta-RL, which facilitates knowledge sharing at the cluster level. To facilitate exploration, we also develop a dedicated cluster-level exploratory policy to discover task clusters via divide-and-conquer. The knowledge from the discovered clusters helps to narrow the search space of task-specific policy learning, leading to more sample-efficient policy adaptation. We evaluate the proposed method on environments with parametric clusters (e.g., rewards and state dynamics in the MuJoCo suite) and non-parametric clusters (e.g., control skills in the Meta-World suite). The results demonstrate strong advantages of our solution against a set of representative meta-RL methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de0a8f62-0659-4685-bd3e-11893093ddcdCited by top-tier papers5
- Mixture-of-Experts Meets In-Context Reinforcement LearningWenhao Wu, Fuhong Liu, Haoru Li, Zican Hu et al.NeurIPS 2025 · 15 citations
- Multi-Objective Intrinsic Reward Learning for Conversational Recommender SystemsZhendong Chu, Nan Wang, Hongning WangNeurIPS 2023 · 5 citations
- CERTAIN: Context Uncertainty-aware One-Shot Adaptation for Context-based Offline Meta Reinforcement LearningHongtu Zhou, Ruiling Yang, Yakun Zhu, Haoqi Zhao et al.ICML 2025
- Regret-Guided Search Control for Efficient Learning in AlphaZeroYun-Jui Tsai, Wei-Yu Chen, Yan-Ru Ju, Yu-Hung Chang et al.ICLR 2026
- Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic SystemsSaptarshi Nath, Christos Peridis, Eseoghene Benjamin, Xinran Liu et al.AAAI 2026
Builds on4
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze et al.ICLR 2020 · 315 citations
- Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without SacrificesEvan Zheran Liu, Aditi Raghunathan, Percy Liang, Chelsea FinnICML 2021 · 80 citations
- Skill-based Meta-Reinforcement LearningTaewook Nam, Shao-Hua Sun, Karl Pertsch, Sung Ju Hwang et al.ICLR 2022 · 55 citations
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen et al.ICML 2021 · 33 citations
Related papers
- Mixture of Meta-Policies for Cross-Environment Meta-Reinforcement LearningXinyu Liu, Qingyu Zeng, Chenwei Tang, Jiancheng LvKDD 2026
- MetaCARD: Meta-Reinforcement Learning with Task Uncertainty Feedback via Decoupled Context-Aware Reward and Dynamics ComponentsMin Wang, Xin Li, Leiji Zhang, Mingzhong WangAAAI 2024 · 6 citations
- Learning and Planning Multi-Agent Tasks via an MoE-based World ModelZijie Zhao, Zhongyue Zhao, Kaixuan Xu, Yuqian Fu et al.NeurIPS 2025 · 12 citations
- Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement LearningJinmin He, Kai Li, Yifan Zang, Haobo Fu et al.ICML 2025
- Learning Generalizable Skills from Offline Multi-Task Data for Multi-Agent CooperationSicong Liu, Yang Shu, Chenjuan Guo, Bin YangICLR 2025
