Entropy Regularized Task Representation Learning for Offline Meta-Reinforcement Learning
Mohammadreza Nakhaeinezhadfard, Aidan Scannell, Joni Pajarinen
摘要
Offline meta-reinforcement learning aims to equip agents with the ability to rapidly adapt to new tasks by training on data from a set of different tasks. Context-based approaches utilize a history of state-action-reward transitions -referred to as the context -to infer representations of the current task, and then condition the agent, i.e., the policy and value function, on the task representations. Intuitively, the better the task representations capture the underlying tasks, the better the agent can generalize to new tasks. Unfortunately, contextbased approaches suffer from distribution mismatch, as the context in the offline data does not match the context at test time, limiting their ability to generalize to the test tasks. This leads to the task representations overfitting to the offline training data. Intuitively, the task representations should be independent of the behavior policy used to collect the offline data. To address this issue, we approximately minimize the mutual information between the distribution over the task representations and behavior policy by maximizing the entropy of behavior policy conditioned on the task representations. We validate our approach in MuJoCo environments, showing that compared to baselines, our task representations more faithfully represent the underlying tasks, leading to outperforming prior methods in both in-distribution and out-of-distribution tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Offline Meta-Reinforcement Learning with Flow-Based Task Inference and Adaptive Correction of Feature OvergeneralizationMin Wang, Xin Li, Mingzhong Wang, Hasnaa BennisAAAI 2026
- Improving Generalization in Offline Meta-Reinforcement Learning via Cross-task ContextsHongcai He, Zetao Zheng, Anjie Zhu, Deqiang Ouyang 等AAAI 2026
它引用的顶会 Paper7
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 等ICML 2020 · 被引用 512 次
- Multi-task Batch Reinforcement Learning with Metric LearningJiachen Li, Quan Vuong, Shuang Liu, Minghua Liu 等NeurIPS 2020 · 被引用 64 次
- FOCAL: Efficient Fully-Offline Meta-Reinforcement Learning via Distance Metric Learning and Behavior RegularizationLanqing Li, Rui Yang, Dijun LuoICLR 2021 · 被引用 64 次
- Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive LearningHaoqi Yuan, Zongqing LuICML 2022 · 被引用 53 次
- Context Shift Reduction for Offline Meta-Reinforcement LearningYunkai Gao, Rui Zhang, Jiaming Guo, Fan Wu 等NeurIPS 2023 · 被引用 30 次
相关 Paper
- CATAL: Causally Disentangled Task Representation Learning for Offline Meta-Reinforcement LearningShan Cong, Chao Yu, Xiangyuan LanAAAI 2026
- Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement LearningLanqing Li, Hai Zhang, Xinyu Zhang, Shatong Zhu 等NeurIPS 2024 · 被引用 24 次
- Scrutinize What We Ignore: Reining In Task Representation Shift Of Context-Based Offline Meta Reinforcement LearningHai Zhang, Boyuan Zheng, Tianying Ji, Jinhang Liu 等ICLR 2025
- Meta-Reinforcement Learning Based on Self-Supervised Task Representation LearningMingyang Wang, Zhenshan Bing, Xiangtong Yao, Shuai Wang 等AAAI 2023 · 被引用 22 次
- Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement LearningFuyuan Qian, Menglong Zhang, Song Wang, Quanying LiuICML 2026
