Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
Brahma S. Pavse, Matthew Zurek, Yudong Chen, Qiaomin Xie, Josiah P. Hanna
Abstract
In many reinforcement learning (RL) applications, we want policies that reach desired states and then keep the controlled system within an acceptable region around the desired states over an indefinite period of time. This latter objective is called stability and is especially important when the state space is unbounded, such that the states can be arbitrarily far from each other and the agent can drift far away from the desired states. For example, in stochastic queuing networks, where queues of waiting jobs can grow without bound, the desired state is all-zero queue lengths. Here, a stable policy ensures queue lengths are finite while an optimal policy minimizes queue lengths. Since an optimal policy is also stable, one would expect that RL algorithms would implicitly give us stable policies. However, in this work, we find that deep RL algorithms that directly minimize the distance to the desired state during online training often result in unstable policies, i.e., policies that drift far away from the desired state. We attribute this instability to poor credit-assignment for destabilizing actions. We then introduce an approach based on two ideas: 1) a Lyapunov-based cost-shaping technique and 2) state transformations to the unbounded state space. We conduct an empirical study on various queueing networks and traffic signal control problems and find that our approach performs competitively against strong baselines with knowledge of the transition dynamics. Our code is available here: https://github.com/Badger-RL/STOP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5657eee1-84cb-4aae-bd3f-e1d82ee8b135Builds on11
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- How Neural Networks Extrapolate: From Feedforward to Graph Neural NetworksKeyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du et al.ICLR 2021 · 364 citations
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah et al.ICLR 2020 · 202 citations
- Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision ProcessesChen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo, Hiteshi Sharma et al.ICML 2020 · 120 citations
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 82 citations
Related papers
- Lyapunov Density Models: Constraining Distribution Shift in Learning-Based ControlKatie Kang, Paula Gradu, Jason J. Choi, Michael Janner et al.ICML 2022 · 39 citations
- Certifying Stability of Reinforcement Learning Policies using Generalized Lyapunov FunctionsKehan Long, Jorge Cortés, Nikolay AtanasovNeurIPS 2025 · 8 citations
- Hierarchically and Cooperatively Learning Traffic Signal ControlBingyu Xu, Yaowei Wang, Zhaozhi Wang, Huizhu Jia et al.AAAI 2021 · 88 citations
- Drift Plus Optimistic Penalty - A Learning Framework for Stochastic Network OptimizationSathwik Chadaga, Eytan H. ModianoINFOCOM 2025 · 1 citation
- Almost Surely Stable Deep DynamicsNathan P. Lawrence, Philip D. Loewen, Michael G. Forbes, Johan U. Backström et al.NeurIPS 2020 · 28 citations
