Continual Learning In Environments With Polynomial Mixing Times
Matthew Riemer, Sharath Chandra Raparthy, Ignacio Cases, Gopeshh Subbaraj, Maximilian Puelma Touzel, Irina Rish
Abstract
The mixing time of the Markov chain induced by a policy limits performance in real-world continual learning scenarios. Yet, the effect of mixing times on learning in continual reinforcement learning (RL) remains underexplored. In this paper, we characterize problems that are of long-term interest to the development of continual RL, which we call scalable MDPs, through the lens of mixing times. In particular, we theoretically establish that scalable MDPs have mixing times that scale polynomially with the size of the problem. We go on to demonstrate that polynomial mixing times present significant difficulties for existing approaches, which suffer from myopic bias and stale bootstrapped estimates. To validate our theory, we study the empirical scaling behavior of mixing times with respect to the number of tasks and task duration for high performing policies deployed across multiple Atari games. Our analysis demonstrates both that polynomial mixing times do emerge in practice and how their existence may lead to unstable learning behavior like catastrophic forgetting in continual learning settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ddb3843-0abd-46b6-bbbb-a551ea3488d4Cited by top-tier papers4
- A Definition of Continual Reinforcement LearningDavid Abel, André Barreto, Benjamin Van Roy, Doina Precup et al.NeurIPS 2023 · 167 citations
- Beyond Exponentially Fast Mixing in Average-Reward Reinforcement Learning via Multi-Level Monte Carlo Actor-CriticWesley A. Suttle, Amrit S. Bedi, Bhrij Patel, Brian M. Sadler et al.ICML 2023 · 24 citations
- : On-Device Real-Time Deep Reinforcement Learning for Autonomous RoboticsZexin Li, Aritra Samanta, Yufei Li, Andrea Soltoggio et al.RTSS 2023 · 9 citations
- Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time OraclesBhrij Patel, Wesley A. Suttle, Alec Koppel, Vaneet Aggarwal et al.ICML 2024 · 4 citations
Builds on7
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 82 citations
- Doubly Robust Bias Reduction in Infinite Horizon Off-Policy EstimationZiyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou et al.ICLR 2020 · 72 citations
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White et al.ICML 2020 · 72 citations
- A Policy Gradient Algorithm for Learning to Learn in Multiagent Reinforcement LearningDong-Ki Kim, Miao Liu, Matthew Riemer, Chuangchuang Sun et al.ICML 2021 · 66 citations
- Influencing Long-Term Behavior in Multiagent Reinforcement LearningDong-Ki Kim, Matthew Riemer, Miao Liu, Jakob N. Foerster et al.NeurIPS 2022 · 29 citations
Related papers
- Optimal Continual Learning has Perfect Memory and is NP-hardJeremias Knoblauch, Hisham Husain, Tom DietheICML 2020 · 116 citations
- Building a Subspace of Policies for Scalable Continual LearningJean-Baptiste Gaya, Thang Doan, Lucas Caccia, Laure Soulier et al.ICLR 2023 · 3 citations
- Temporal-Difference Variational Continual LearningLuckeciano Carvalho Melo, Alessandro Abate, Yarin GalNeurIPS 2025 · 1 citation
- Continual World: A Robotic Benchmark For Continual Reinforcement LearningMaciej Wolczyk, Michal Zajac, Razvan Pascanu, Lukasz Kucinski et al.NeurIPS 2021 · 152 citations
- Continual Predictive Learning from VideosGeng Chen, Wendong Zhang, Han Lu, Siyu Gao et al.CVPR 2022 · 5 citations
