Skill-aware Mutual Information Optimisation for Zero-shot Generalisation in Reinforcement Learning
Xuehui Yu, Mhairi Dunion, Xin Li, Stefano V. Albrecht
摘要
Meta-Reinforcement Learning (Meta-RL) agents can struggle to operate across tasks with varying environmental features that require different optimal skills (i.e., different modes of behaviour). Using context encoders based on contrastive learning to enhance the generalisability of Meta-RL agents is now widely studied but faces challenges such as the requirement for a large sample size, also referred to as the log-K curse. To improve RL generalisation to different tasks, we first introduce Skill-aware Mutual Information (SaMI), an optimisation objective that aids in distinguishing context embeddings according to skills, thereby equipping RL agents with the ability to identify and execute different skills across tasks. We then propose Skill-aware Noise Contrastive Estimation (SaNCE), a K-sample estimator used to optimise the SaMI objective. We provide a framework for equipping an RL agent with SaNCE in practice and conduct experimental validation on modified MuJoCo and Panda-gym benchmarks. We empirically find that RL agents that learn by maximising SaMI achieve substantially improved zero-shot generalisation to unseen tasks. Additionally, the context encoder trained with SaNCE demonstrates greater robustness to a reduction in the number of available samples, thus possessing the potential to overcome the log-K curse.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Self-Improving Skill Learning for Robust Skill-based Meta-Reinforcement LearningSeungyul Han, Sanghyeon Lee, Sangjun Bae, Yisak ParkICLR 2026 · 被引用 5 次
- Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement LearningFuyuan Qian, Menglong Zhang, Song Wang, Quanying LiuICML 2026
- Conflict-Aware Additive Guidance for Flow Models under Compositional RewardsXuehui Yu, Fucheng Cai, Meiyi Wang, Xiaopeng Fan 等ICML 2026
它引用的顶会 Paper16
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 331 次
- Understanding the Limitations of Variational Mutual Information EstimatorsJiaming Song, Stefano ErmonICLR 2020 · 被引用 243 次
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement LearningKimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee 等ICML 2020 · 被引用 158 次
相关 Paper
- Behavior Contrastive Learning for Unsupervised Skill DiscoveryRushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li 等ICML 2023 · 被引用 34 次
- Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive LearningHaoqi Yuan, Zongqing LuICML 2022 · 被引用 53 次
- Generalizable Task Representation Learning for Offline Meta-Reinforcement Learning with Data LimitationsRenzhe Zhou, Chenxiao Gao, Zongzhang Zhang, Yang YuAAAI 2024 · 被引用 16 次
- Towards Effective Context for Meta-Reinforcement Learning: an Approach based on Contrastive LearningHaotian Fu, Hongyao Tang, Jianye Hao, Chen Chen 等AAAI 2021 · 被引用 61 次
- Decoupling Meta-Reinforcement Learning with Gaussian Task Contexts and SkillsHongcai He, Anjie Zhu, Shuang Liang, Feiyu Chen 等AAAI 2024 · 被引用 6 次
