Lune

NeurIPS2023顶会

Tempo Adaptation in Non-stationary Reinforcement Learning

Hyunin Lee, Yuhao Ding, Jongmin Lee, Ming Jin, Javad Lavaei, Somayeh Sojoudi

2023年份
6被引次数
3顶会引用

摘要

We first raise and tackle a ``time synchronization'' issue between the agent and the environment in non-stationary reinforcement learning (RL), a crucial factor hindering its real-world applications. In reality, environmental changes occur over wall-clock time (tt) rather than episode progress (kk), where wall-clock time signifies the actual elapsed time within the fixed duration t∈[0,T]t \in [0, T]. In existing works, at episode kk, the agent rolls a trajectory and trains a policy before transitioning to episode k+1k+1. In the context of the time-desynchronized environment, however, the agent at time tkt_{k} allocates Δt\Delta t for trajectory generation and training, subsequently moves to the next episode at tk+1=tk+Δtt_{k+1}=t_{k}+\Delta t. Despite a fixed total number of episodes (KK), the agent accumulates different trajectories influenced by the choice of interaction times (t1,t2,...,tKt_1,t_2,...,t_K), significantly impacting the suboptimality gap of the policy. We propose a Proactively Synchronizing Tempo (ProST\texttt{ProST}) framework that computes a suboptimal sequence t1,t2,...,tKt_1,t_2,...,t_K (= t1:Kt_{1:K}) by minimizing an upper bound on its performance measure, i.e., the dynamic regret. Our main contribution is that we show that a suboptimal t1:Kt_{1:K} trades-off between the policy training time (agent tempo) and how fast the environment changes (environment tempo). Theoretically, this work develops a suboptimal t1:Kt_{1:K} as a function of the degree of the environment's non-stationarity while also achieving a sublinear dynamic regret. Our experimental evaluation on various high-dimensional non-stationary environments shows that the ProST\texttt{ProST} framework achieves a higher online return at suboptimal t1:Kt_{1:K} than the existing methods.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖