Tempo Adaptation in Non-stationary Reinforcement Learning
Hyunin Lee, Yuhao Ding, Jongmin Lee, Ming Jin, Javad Lavaei, Somayeh Sojoudi
Abstract
We first raise and tackle a ``time synchronization'' issue between the agent and the environment in non-stationary reinforcement learning (RL), a crucial factor hindering its real-world applications. In reality, environmental changes occur over wall-clock time () rather than episode progress (), where wall-clock time signifies the actual elapsed time within the fixed duration . In existing works, at episode , the agent rolls a trajectory and trains a policy before transitioning to episode . In the context of the time-desynchronized environment, however, the agent at time allocates for trajectory generation and training, subsequently moves to the next episode at . Despite a fixed total number of episodes (), the agent accumulates different trajectories influenced by the choice of interaction times (), significantly impacting the suboptimality gap of the policy. We propose a Proactively Synchronizing Tempo () framework that computes a suboptimal sequence (= ) by minimizing an upper bound on its performance measure, i.e., the dynamic regret. Our main contribution is that we show that a suboptimal trades-off between the policy training time (agent tempo) and how fast the environment changes (environment tempo). Theoretically, this work develops a suboptimal as a function of the degree of the environment's non-stationarity while also achieving a sublinear dynamic regret. Our experimental evaluation on various high-dimensional non-stationary environments shows that the framework achieves a higher online return at suboptimal than the existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Pausing Policy Learning in Non-stationary Reinforcement LearningHyunin Lee, Ming Jin, Javad Lavaei, Somayeh SojoudiICML 2024 · 4 citations
- Online Reinforcement Learning in Non-Stationary Context-Driven EnvironmentsPouya Hamadanian, Arash Nasr-Esfahany, Malte Schwarzkopf, Siddhartha Sen et al.ICLR 2025
- Forecasting in Offline Reinforcement Learning for Non-stationary EnvironmentsSuzan Ece Ada, Georg Martius, Emre Ugur, Erhan OztopNeurIPS 2025
Builds on11
- OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationJongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau et al.ICML 2021 · 137 citations
- Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) OptimismWang Chi Cheung, David Simchi-Levi, Ruihao ZhuICML 2020 · 114 citations
- RL for Latent MDPs: Regret Guarantees and a Lower BoundJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 91 citations
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess et al.ICLR 2022 · 84 citations
- AdaRL: What, Where, and How to Adapt in Transfer Reinforcement LearningBiwei Huang, Fan Feng, Chaochao Lu, Sara Magliacane et al.ICLR 2022 · 75 citations
Related papers
- Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement LearningSid Bharthulwar, Stone Tao, Hao SuNeurIPS 2025 · 6 citations
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White et al.ICML 2020 · 72 citations
- Provably Efficient Algorithm for Nonstationary Low-Rank MDPsYuan Cheng, Jing Yang, Yingbin LiangNeurIPS 2023 · 2 citations
- Dynamic Regret of Policy Optimization in Non-Stationary EnvironmentsYingjie Fei, Zhuoran Yang, Zhaoran Wang, Qiaomin XieNeurIPS 2020 · 73 citations
- Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and ConstraintsYuhao Ding, Javad LavaeiAAAI 2023 · 32 citations
