Reinforcement Learning in Presence of Discrete Markovian Context Evolution
Hang Ren, Aivar Sootla, Taher Jafferjee, Junxiao Shen, Jun Wang, Haitham Bou-Ammar
Abstract
We consider a context-dependent Reinforcement Learning (RL) setting, which is characterized by: a) an unknown finite number of not directly observable contexts; b) abrupt (discontinuous) context changes occurring during an episode; and c) Markovian context evolution. We argue that this challenging case is often met in applications and we tackle it using a Bayesian approach and variational inference. We adapt a sticky Hierarchical Dirichlet Process (HDP) prior for model learning, which is arguably best-suited for Markov process modeling. We then derive a context distillation procedure, which identifies and removes spurious contexts in an unsupervised fashion. We argue that the combination of these two components allows to infer the number of contexts from data thus dealing with the context cardinality assumption. We then find the representation of the optimal policy enabling efficient policy learning using off-the-shelf RL algorithms. Finally, we demonstrate empirically (using gym environments cart-pole swing-up, drone, intersection) that our approach succeeds where state-of-the-art methods of other frameworks fail and elaborate on the reasons for such failures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bcc2d6f-997d-470f-82de-ce436d02a1e6Cited by top-tier papers4
- An Adaptive Deep RL Method for Non-Stationary Environments with Piecewise Stable ContextXiaoyu Chen, Xiangming Zhu, Yufeng Zheng, Pushi Zhang et al.NeurIPS 2022 · 24 citations
- Building a Subspace of Policies for Scalable Continual LearningJean-Baptiste Gaya, Thang Doan, Lucas Caccia, Laure Soulier et al.ICLR 2023 · 3 citations
- Online Reinforcement Learning in Non-Stationary Context-Driven EnvironmentsPouya Hamadanian, Arash Nasr-Esfahany, Malte Schwarzkopf, Siddhartha Sen et al.ICLR 2025
- Wavelet Predictive Representations for Non-Stationary Reinforcement LearningMin Wang, Xin Li, Ye He, Yao-Hui Li et al.ICLR 2026
Builds on6
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze et al.ICLR 2020 · 315 citations
- Invariant Causal Prediction for Block MDPsAmy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos et al.ICML 2020 · 153 citations
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White et al.ICML 2020 · 72 citations
- Task-Agnostic Online Reinforcement Learning with an Infinite Mixture of Gaussian ProcessesMengdi Xu, Wenhao Ding, Jiacheng Zhu, Zuxin Liu et al.NeurIPS 2020 · 39 citations
Related papers
- Behavior-agnostic Task Inference for Robust Offline In-context Reinforcement LearningLong Ma, Fangwei Zhong, Yizhou WangICML 2025
- MetaCARD: Meta-Reinforcement Learning with Task Uncertainty Feedback via Decoupled Context-Aware Reward and Dynamics ComponentsMin Wang, Xin Li, Leiji Zhang, Mingzhong WangAAAI 2024 · 6 citations
- Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot GeneralizationFrank Röder, Jan Benad, Manfred Eppe, Pradeep Kr. BanerjeeNeurIPS 2025 · 9 citations
- Distilling Reinforcement Learning Algorithms for In-Context Model-Based PlanningJaehyeon Son, Soochan Lee, Gunhee KimICLR 2025
- Behaviour DistillationAndrei Lupu, Chris Lu, Jarek Liesen, Robert Tjarko Lange et al.ICLR 2024 · 8 citations
