Mixture of Meta-Policies for Cross-Environment Meta-Reinforcement Learning
Xinyu Liu, Qingyu Zeng, Chenwei Tang, Jiancheng Lv
摘要
Meta-Reinforcement Learning (Meta-RL) aims to enable agents to rapidly adapt to new tasks by leveraging prior experience from related ones. However, existing approaches typically assume consistent state and action spaces across tasks, limiting the ability of Meta-RL methods to generalize to environments with diverse morphologies, dynamics, and reward structures, common in real-world applications. To address this, we propose Mixture of Meta-Policies (MoMP), a modular and scalable Meta-RL framework designed for effective knowledge transfer across structurally heterogeneous environments. MoMP represents policies as sparse combinations of shared submodules, where each submodule is itself an attention-based policy block. During meta-training, MoMP learns to specialize and reuse these attention modules by optimizing their task-conditioned activation across multiple environments. Once trained, the shared module can be plugged into new agents to accelerate learning in novel tasks. Sparse activation enables targeted reuse of prior knowledge while mitigating interference, improving both adaptation speed and long-term retention. Experiments across diverse MuJoCo agents show that MoMP significantly outperforms strong meta-RL baselines, highlighting its effectiveness in cross-environment generalization and efficient meta-policy reuse.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- HMRL: Hyper-Meta Learning for Sparse Reward Reinforcement Learning ProblemYun Hua, Xiangfeng Wang, Bo Jin, Wenhao Li 等KDD 2021 · 被引用 6 次
- Fast And Slow Learning Of Recurrent Independent MechanismsKanika Madan, Nan Rosemary Ke, Anirudh Goyal, Bernhard Schölkopf 等ICLR 2021 · 被引用 41 次
- Meta-Reinforcement Learning via Exploratory Task ClusteringZhendong Chu, Renqin Cai, Hongning WangAAAI 2024 · 被引用 12 次
- Learning and Planning Multi-Agent Tasks via an MoE-based World ModelZijie Zhao, Zhongyue Zhao, Kaixuan Xu, Yuqian Fu 等NeurIPS 2025 · 被引用 12 次
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen 等ICML 2021 · 被引用 33 次
