Lune

KDD2026顶会

Mixture of Meta-Policies for Cross-Environment Meta-Reinforcement Learning

Xinyu Liu, Qingyu Zeng, Chenwei Tang, Jiancheng Lv

2026年份

摘要

Meta-Reinforcement Learning (Meta-RL) aims to enable agents to rapidly adapt to new tasks by leveraging prior experience from related ones. However, existing approaches typically assume consistent state and action spaces across tasks, limiting the ability of Meta-RL methods to generalize to environments with diverse morphologies, dynamics, and reward structures, common in real-world applications. To address this, we propose Mixture of Meta-Policies (MoMP), a modular and scalable Meta-RL framework designed for effective knowledge transfer across structurally heterogeneous environments. MoMP represents policies as sparse combinations of shared submodules, where each submodule is itself an attention-based policy block. During meta-training, MoMP learns to specialize and reuse these attention modules by optimizing their task-conditioned activation across multiple environments. Once trained, the shared module can be plugged into new agents to accelerate learning in novel tasks. Sparse activation enables targeted reuse of prior knowledge while mitigating interference, improving both adaptation speed and long-term retention. Experiments across diverse MuJoCo agents show that MoMP significantly outperforms strong meta-RL baselines, highlighting its effectiveness in cross-environment generalization and efficient meta-policy reuse.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get be09f91b-654c-4f5a-94de-6621bf819acd

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖