Lune

KDD2026Top-tier venue

Mixture of Meta-Policies for Cross-Environment Meta-Reinforcement Learning

Xinyu Liu, Qingyu Zeng, Chenwei Tang, Jiancheng Lv

2026Year

Abstract

Meta-Reinforcement Learning (Meta-RL) aims to enable agents to rapidly adapt to new tasks by leveraging prior experience from related ones. However, existing approaches typically assume consistent state and action spaces across tasks, limiting the ability of Meta-RL methods to generalize to environments with diverse morphologies, dynamics, and reward structures, common in real-world applications. To address this, we propose Mixture of Meta-Policies (MoMP), a modular and scalable Meta-RL framework designed for effective knowledge transfer across structurally heterogeneous environments. MoMP represents policies as sparse combinations of shared submodules, where each submodule is itself an attention-based policy block. During meta-training, MoMP learns to specialize and reuse these attention modules by optimizing their task-conditioned activation across multiple environments. Once trained, the shared module can be plugged into new agents to accelerate learning in novel tasks. Sparse activation enables targeted reuse of prior knowledge while mitigating interference, improving both adaptation speed and long-term retention. Experiments across diverse MuJoCo agents show that MoMP significantly outperforms strong meta-RL baselines, highlighting its effectiveness in cross-environment generalization and efficient meta-policy reuse.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get be09f91b-654c-4f5a-94de-6621bf819acd

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines