Mixture of Meta-Policies for Cross-Environment Meta-Reinforcement Learning
Xinyu Liu, Qingyu Zeng, Chenwei Tang, Jiancheng Lv
Abstract
Meta-Reinforcement Learning (Meta-RL) aims to enable agents to rapidly adapt to new tasks by leveraging prior experience from related ones. However, existing approaches typically assume consistent state and action spaces across tasks, limiting the ability of Meta-RL methods to generalize to environments with diverse morphologies, dynamics, and reward structures, common in real-world applications. To address this, we propose Mixture of Meta-Policies (MoMP), a modular and scalable Meta-RL framework designed for effective knowledge transfer across structurally heterogeneous environments. MoMP represents policies as sparse combinations of shared submodules, where each submodule is itself an attention-based policy block. During meta-training, MoMP learns to specialize and reuse these attention modules by optimizing their task-conditioned activation across multiple environments. Once trained, the shared module can be plugged into new agents to accelerate learning in novel tasks. Sparse activation enables targeted reuse of prior knowledge while mitigating interference, improving both adaptation speed and long-term retention. Experiments across diverse MuJoCo agents show that MoMP significantly outperforms strong meta-RL baselines, highlighting its effectiveness in cross-environment generalization and efficient meta-policy reuse.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get be09f91b-654c-4f5a-94de-6621bf819acdRelated papers
- HMRL: Hyper-Meta Learning for Sparse Reward Reinforcement Learning ProblemYun Hua, Xiangfeng Wang, Bo Jin, Wenhao Li et al.KDD 2021 · 6 citations
- Fast And Slow Learning Of Recurrent Independent MechanismsKanika Madan, Nan Rosemary Ke, Anirudh Goyal, Bernhard Schölkopf et al.ICLR 2021 · 41 citations
- Meta-Reinforcement Learning via Exploratory Task ClusteringZhendong Chu, Renqin Cai, Hongning WangAAAI 2024 · 12 citations
- Learning and Planning Multi-Agent Tasks via an MoE-based World ModelZijie Zhao, Zhongyue Zhao, Kaixuan Xu, Yuqian Fu et al.NeurIPS 2025 · 12 citations
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen et al.ICML 2021 · 33 citations
