Tempo Adaptation in Non-stationary Reinforcement Learning
Hyunin Lee, Yuhao Ding, Jongmin Lee, Ming Jin, Javad Lavaei, Somayeh Sojoudi
摘要
We first raise and tackle a ``time synchronization'' issue between the agent and the environment in non-stationary reinforcement learning (RL), a crucial factor hindering its real-world applications. In reality, environmental changes occur over wall-clock time () rather than episode progress (), where wall-clock time signifies the actual elapsed time within the fixed duration . In existing works, at episode , the agent rolls a trajectory and trains a policy before transitioning to episode . In the context of the time-desynchronized environment, however, the agent at time allocates for trajectory generation and training, subsequently moves to the next episode at . Despite a fixed total number of episodes (), the agent accumulates different trajectories influenced by the choice of interaction times (), significantly impacting the suboptimality gap of the policy. We propose a Proactively Synchronizing Tempo () framework that computes a suboptimal sequence (= ) by minimizing an upper bound on its performance measure, i.e., the dynamic regret. Our main contribution is that we show that a suboptimal trades-off between the policy training time (agent tempo) and how fast the environment changes (environment tempo). Theoretically, this work develops a suboptimal as a function of the degree of the environment's non-stationarity while also achieving a sublinear dynamic regret. Our experimental evaluation on various high-dimensional non-stationary environments shows that the framework achieves a higher online return at suboptimal than the existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Pausing Policy Learning in Non-stationary Reinforcement LearningHyunin Lee, Ming Jin, Javad Lavaei, Somayeh SojoudiICML 2024 · 被引用 4 次
- Online Reinforcement Learning in Non-Stationary Context-Driven EnvironmentsPouya Hamadanian, Arash Nasr-Esfahany, Malte Schwarzkopf, Siddhartha Sen 等ICLR 2025
- Forecasting in Offline Reinforcement Learning for Non-stationary EnvironmentsSuzan Ece Ada, Georg Martius, Emre Ugur, Erhan OztopNeurIPS 2025
它引用的顶会 Paper11
- OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationJongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau 等ICML 2021 · 被引用 137 次
- Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) OptimismWang Chi Cheung, David Simchi-Levi, Ruihao ZhuICML 2020 · 被引用 114 次
- RL for Latent MDPs: Regret Guarantees and a Lower BoundJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 被引用 91 次
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess 等ICLR 2022 · 被引用 84 次
- AdaRL: What, Where, and How to Adapt in Transfer Reinforcement LearningBiwei Huang, Fan Feng, Chaochao Lu, Sara Magliacane 等ICLR 2022 · 被引用 75 次
相关 Paper
- Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement LearningSid Bharthulwar, Stone Tao, Hao SuNeurIPS 2025 · 被引用 6 次
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White 等ICML 2020 · 被引用 72 次
- Provably Efficient Algorithm for Nonstationary Low-Rank MDPsYuan Cheng, Jing Yang, Yingbin LiangNeurIPS 2023 · 被引用 2 次
- Dynamic Regret of Policy Optimization in Non-Stationary EnvironmentsYingjie Fei, Zhuoran Yang, Zhaoran Wang, Qiaomin XieNeurIPS 2020 · 被引用 73 次
- Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and ConstraintsYuhao Ding, Javad LavaeiAAAI 2023 · 被引用 32 次
