Meta-Reinforcement Learning Robust to Distributional Shift Via Performing Lifelong In-Context Learning
Tengye Xu, Zihao Li, Qinyuan Ren
Abstract
A key challenge in Meta-Reinforcement Learning (meta-RL) is the task distribution shift, since the generalization ability of most current meta-RL methods is limited to tasks sampled from the training distribution. In this paper, we propose Posterior Sampling Bayesian Lifelong In-Context Reinforcement Learning (PSBL), which is robust to task distribution shift. PSBL meta-trains a variant of transformer to directly perform amortized inference about the Predictive Posterior Distribution (PPD) of the optimal policy. Once trained, the network can infer the PPD online with frozen parameters. The agent then samples actions from the approximate PPD to perform online exploration, which progressively reduces uncertainty and enhances performance in the interaction with the environment. This property is known as in-context learning. Experimental results demonstrate that PSBL significantly outperforms standard Meta RL methods both in tasks with sparse rewards and dense rewards when the test task distribution is strictly shifted from the training distribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a80dd61-cbe4-4c42-8f8e-3d8932151d8cCited by top-tier papers8
- Reward Is Enough: LLMs Are In-Context Reinforcement LearnersKefan Song, Amir Moeini, Peng Wang, Lei Gong et al.ICLR 2026 · 42 citations
- Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement LearningHaozhe Ma, Zhengding Luo, Thanh Vinh Vo, Kuankuan Sima et al.NeurIPS 2025 · 9 citations
- Towards Provable Emergence of In-Context Reinforcement LearningJiuqi Wang, Rohan Chandra, Shangtong ZhangNeurIPS 2025 · 5 citations
- A Bayesian Fast-Slow Framework to Mitigate Interference in Non-Stationary Reinforcement LearningYihuan Mao, Chongjie ZhangNeurIPS 2025 · 1 citation
- Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution TasksJeongmo Kim, Yisak Park, Minung Kim, Seungyul HanICML 2025
Builds on13
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka et al.ICLR 2022 · 287 citations
- Stabilizing Transformer Training by Preventing Attention Entropy CollapseShuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge et al.ICML 2023 · 153 citations
- Observational Overfitting in Reinforcement LearningXingyou Song, Yiding Jiang, Stephen Tu, Yilun Du et al.ICLR 2020 · 148 citations
- What learning algorithm is in-context learning? Investigations with linear modelsEkin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma et al.ICLR 2023 · 85 citations
Related papers
- Multi-Task Bayesian In-Context LearningQingyang Zhu, Eric Oermann, Kyunghyun ChoICML 2026
- Supervised Pretraining Can Learn In-Context Reinforcement LearningJonathan Lee, Annie Xie, Aldo Pacchiano, Yash Chandak et al.NeurIPS 2023 · 170 citations
- In-context Exploration-Exploitation for Reinforcement LearningZhenwen Dai, Federico Tomasi, Sina GhiassianICLR 2024 · 14 citations
- Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement LearningFuyuan Qian, Menglong Zhang, Song Wang, Quanying LiuICML 2026
- Behavior-agnostic Task Inference for Robust Offline In-context Reinforcement LearningLong Ma, Fangwei Zhong, Yizhou WangICML 2025
