Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
Brahma S. Pavse, Matthew Zurek, Yudong Chen, Qiaomin Xie, Josiah P. Hanna
摘要
In many reinforcement learning (RL) applications, we want policies that reach desired states and then keep the controlled system within an acceptable region around the desired states over an indefinite period of time. This latter objective is called stability and is especially important when the state space is unbounded, such that the states can be arbitrarily far from each other and the agent can drift far away from the desired states. For example, in stochastic queuing networks, where queues of waiting jobs can grow without bound, the desired state is all-zero queue lengths. Here, a stable policy ensures queue lengths are finite while an optimal policy minimizes queue lengths. Since an optimal policy is also stable, one would expect that RL algorithms would implicitly give us stable policies. However, in this work, we find that deep RL algorithms that directly minimize the distance to the desired state during online training often result in unstable policies, i.e., policies that drift far away from the desired state. We attribute this instability to poor credit-assignment for destabilizing actions. We then introduce an approach based on two ideas: 1) a Lyapunov-based cost-shaping technique and 2) state transformations to the unbounded state space. We conduct an empirical study on various queueing networks and traffic signal control problems and find that our approach performs competitively against strong baselines with knowledge of the transition dynamics. Our code is available here: https://github.com/Badger-RL/STOP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- How Neural Networks Extrapolate: From Feedforward to Graph Neural NetworksKeyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du 等ICLR 2021 · 被引用 364 次
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah 等ICLR 2020 · 被引用 202 次
- Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision ProcessesChen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo, Hiteshi Sharma 等ICML 2020 · 被引用 120 次
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 被引用 82 次
相关 Paper
- Lyapunov Density Models: Constraining Distribution Shift in Learning-Based ControlKatie Kang, Paula Gradu, Jason J. Choi, Michael Janner 等ICML 2022 · 被引用 39 次
- Certifying Stability of Reinforcement Learning Policies using Generalized Lyapunov FunctionsKehan Long, Jorge Cortés, Nikolay AtanasovNeurIPS 2025 · 被引用 8 次
- Hierarchically and Cooperatively Learning Traffic Signal ControlBingyu Xu, Yaowei Wang, Zhaozhi Wang, Huizhu Jia 等AAAI 2021 · 被引用 88 次
- Drift Plus Optimistic Penalty - A Learning Framework for Stochastic Network OptimizationSathwik Chadaga, Eytan H. ModianoINFOCOM 2025 · 被引用 1 次
- Almost Surely Stable Deep DynamicsNathan P. Lawrence, Philip D. Loewen, Michael G. Forbes, Johan U. Backström 等NeurIPS 2020 · 被引用 28 次
