Skill-aware Mutual Information Optimisation for Zero-shot Generalisation in Reinforcement Learning
Xuehui Yu, Mhairi Dunion, Xin Li, Stefano V. Albrecht
Abstract
Meta-Reinforcement Learning (Meta-RL) agents can struggle to operate across tasks with varying environmental features that require different optimal skills (i.e., different modes of behaviour). Using context encoders based on contrastive learning to enhance the generalisability of Meta-RL agents is now widely studied but faces challenges such as the requirement for a large sample size, also referred to as the log-K curse. To improve RL generalisation to different tasks, we first introduce Skill-aware Mutual Information (SaMI), an optimisation objective that aids in distinguishing context embeddings according to skills, thereby equipping RL agents with the ability to identify and execute different skills across tasks. We then propose Skill-aware Noise Contrastive Estimation (SaNCE), a K-sample estimator used to optimise the SaMI objective. We provide a framework for equipping an RL agent with SaNCE in practice and conduct experimental validation on modified MuJoCo and Panda-gym benchmarks. We empirically find that RL agents that learn by maximising SaMI achieve substantially improved zero-shot generalisation to unseen tasks. Additionally, the context encoder trained with SaNCE demonstrates greater robustness to a reduction in the number of available samples, thus possessing the potential to overcome the log-K curse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e7197f4-89a0-4051-b65a-2e06fcffbedcCited by top-tier papers3
- Self-Improving Skill Learning for Robust Skill-based Meta-Reinforcement LearningSeungyul Han, Sanghyeon Lee, Sangjun Bae, Yisak ParkICLR 2026 · 5 citations
- Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement LearningFuyuan Qian, Menglong Zhang, Song Wang, Quanying LiuICML 2026
- Conflict-Aware Additive Guidance for Flow Models under Compositional RewardsXuehui Yu, Fucheng Cai, Meiyi Wang, Xiaopeng Fan et al.ICML 2026
Builds on16
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
- Understanding the Limitations of Variational Mutual Information EstimatorsJiaming Song, Stefano ErmonICLR 2020 · 243 citations
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement LearningKimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee et al.ICML 2020 · 158 citations
Related papers
- Behavior Contrastive Learning for Unsupervised Skill DiscoveryRushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li et al.ICML 2023 · 34 citations
- Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive LearningHaoqi Yuan, Zongqing LuICML 2022 · 53 citations
- Generalizable Task Representation Learning for Offline Meta-Reinforcement Learning with Data LimitationsRenzhe Zhou, Chenxiao Gao, Zongzhang Zhang, Yang YuAAAI 2024 · 16 citations
- Towards Effective Context for Meta-Reinforcement Learning: an Approach based on Contrastive LearningHaotian Fu, Hongyao Tang, Jianye Hao, Chen Chen et al.AAAI 2021 · 61 citations
- Decoupling Meta-Reinforcement Learning with Gaussian Task Contexts and SkillsHongcai He, Anjie Zhu, Shuang Liang, Feiyu Chen et al.AAAI 2024 · 6 citations
