Simultaneously Updating All Persistence Values in Reinforcement Learning
Luca Sabbioni, Luca Al Daire, Lorenzo Bisi, Alberto Maria Metelli, Marcello Restelli
Abstract
In Reinforcement Learning, the performance of learning agents is highly sensitive to the choice of time discretization. Agents acting at high frequencies have the best control opportunities, along with some drawbacks, such as possible inefficient exploration and vanishing of the action advantages. The repetition of the actions, i.e., action persistence, comes into help, as it allows the agent to visit wider regions of the state space and improve the estimation of the action effects. In this work, we derive a novel operator, the All-Persistence Bellman Operator, which allows an effective use of both the low-persistence experience, by decomposition into sub-transition, and the high-persistence experience, thanks to the introduction of a suitable bootstrap procedure. In this way, we employ transitions collected at any time scale to update simultaneously the action values of the considered persistence set. We prove the contraction property of the All-Persistence Bellman Operator and, based on it, we extend classic Q-learning and DQN. After providing a study on the effects of persistence, we experimentally evaluate our approach in both tabular contexts and more challenging frameworks, including some Atari games.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Control Frequency Adaptation via Action Persistence in Batch Reinforcement LearningAlberto Maria Metelli, Flavio Mazzolini, Lorenzo Bisi, Luca Sabbioni et al.ICML 2020 · 43 citations
- TempoRL: Learning When to ActAndré Biedenkapp, Raghu Rajan, Frank Hutter, Marius LindauerICML 2021 · 38 citations
- Time Discretization-Invariant Safe Action Repetition for Policy Gradient MethodsSeohong Park, Jaekyeom Kim, Gunhee KimNeurIPS 2021 · 33 citations
- Locally Persistent Exploration in Continuous Control Tasks with Sparse RewardsSusan Amin, Maziar Gomrokchi, Hossein Aboutalebi, Harsh Satija et al.ICML 2021 · 17 citations
Related papers
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 56 citations
- Bayesian Bellman OperatorsMattie Fellows, Kristian Hartikainen, Shimon WhitesonNeurIPS 2021 · 20 citations
- Reinforcement Learning for Control with Multiple FrequenciesJongmin Lee, Byung-Jun Lee, Kee-Eung KimNeurIPS 2020 · 19 citations
- Parameterized Projected Bellman OperatorThéo Vincent, Alberto Maria Metelli, Boris Belousov, Jan Peters et al.AAAI 2024 · 6 citations
- Exploiting the Replay Memory Before Exploring the Environment: Enhancing Reinforcement Learning Through Empirical MDP IterationHongming Zhang, Chenjun Xiao, Chao Gao, Han Wang et al.NeurIPS 2024 · 7 citations
