Reinforcement Learning for Control with Multiple Frequencies
Jongmin Lee, Byung-Jun Lee, Kee-Eung Kim
Abstract
Many real-world sequential decision problems involve multiple action variables whose control frequencies are different, such that actions take their effects at different periods. While these problems can be formulated with the notion of multiple action persistences in factored-action MDP (FA-MDP), it is non-trivial to solve them efficiently since an action-persistent policy constructed from a stationary policy can be arbitrarily suboptimal, rendering solution methods for the standard FA-MDPs hardly applicable. In this paper, we formalize the problem of multiple control frequencies in RL and provide its efficient solution method. Our proposed method, Action-Persistent Policy Iteration (AP-PI), provides a theoretical guarantee on the convergence to an optimal solution while incurring only a factor of |A| increase in time complexity during policy improvement step, compared to the standard policy iteration for FA-MDPs. Extending this result, we present Action-Persistent Actor-Critic (AP-AC), a scalable RL algorithm for high-dimensional control tasks. In the experiments, we demonstrate that AP-AC significantly outperforms the baselines on several continuous control tasks and a traffic control simulation, which highlights the effectiveness of our method that directly optimizes the periodic non-stationary policy for tasks with multiple control frequencies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3b11050-7295-402c-a748-d67d221f8db5Cited by top-tier papers6
- TempoRL: Learning When to ActAndré Biedenkapp, Raghu Rajan, Frank Hutter, Marius LindauerICML 2021 · 38 citations
- TAAC: Temporally Abstract Actor-Critic for Continuous ControlHaonan Yu, Wei Xu, Haichao ZhangNeurIPS 2021 · 30 citations
- When to Sense and Control? A Time-adaptive Approach for Continuous-Time RLLenart Treven, Bhavya Sukhija, Yarden As, Florian Dörfler et al.NeurIPS 2024 · 10 citations
- Information-Theoretic State Space Model for Multi-View Reinforcement LearningHyeongJoo Hwang, Seokin Seo, Youngsoo Jang, Sungyoon Kim et al.ICML 2023 · 2 citations
- Learning Routines for Effective Off-Policy Reinforcement LearningEdoardo Cetin, Oya ÇeliktutanICML 2021 · 1 citation
Builds on1
Related papers
- Addressing Action Oscillations through Learning Policy InertiaChen Chen, Hongyao Tang, Jianye Hao, Wulong Liu et al.AAAI 2021 · 27 citations
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 150 citations
- Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement LearningTianchen Zhou, Hairi, Haibo Yang, Jia Liu et al.ICML 2024 · 4 citations
- Finite-Time Convergence and Sample Complexity of Multi-Agent Actor-Critic Reinforcement Learning with Average RewardHairi, Jia Liu, Songtao LuICLR 2022 · 21 citations
- Select before Act: Spatially Decoupled Action Repetition for Continuous ControlBuqing Nie, Yangqing Fu, Yue GaoICLR 2025
