Reinforcement Learning for Control with Multiple Frequencies
Jongmin Lee, Byung-Jun Lee, Kee-Eung Kim
摘要
Many real-world sequential decision problems involve multiple action variables whose control frequencies are different, such that actions take their effects at different periods. While these problems can be formulated with the notion of multiple action persistences in factored-action MDP (FA-MDP), it is non-trivial to solve them efficiently since an action-persistent policy constructed from a stationary policy can be arbitrarily suboptimal, rendering solution methods for the standard FA-MDPs hardly applicable. In this paper, we formalize the problem of multiple control frequencies in RL and provide its efficient solution method. Our proposed method, Action-Persistent Policy Iteration (AP-PI), provides a theoretical guarantee on the convergence to an optimal solution while incurring only a factor of |A| increase in time complexity during policy improvement step, compared to the standard policy iteration for FA-MDPs. Extending this result, we present Action-Persistent Actor-Critic (AP-AC), a scalable RL algorithm for high-dimensional control tasks. In the experiments, we demonstrate that AP-AC significantly outperforms the baselines on several continuous control tasks and a traffic control simulation, which highlights the effectiveness of our method that directly optimizes the periodic non-stationary policy for tasks with multiple control frequencies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- TempoRL: Learning When to ActAndré Biedenkapp, Raghu Rajan, Frank Hutter, Marius LindauerICML 2021 · 被引用 38 次
- TAAC: Temporally Abstract Actor-Critic for Continuous ControlHaonan Yu, Wei Xu, Haichao ZhangNeurIPS 2021 · 被引用 30 次
- When to Sense and Control? A Time-adaptive Approach for Continuous-Time RLLenart Treven, Bhavya Sukhija, Yarden As, Florian Dörfler 等NeurIPS 2024 · 被引用 10 次
- Information-Theoretic State Space Model for Multi-View Reinforcement LearningHyeongJoo Hwang, Seokin Seo, Youngsoo Jang, Sungyoon Kim 等ICML 2023 · 被引用 2 次
- Learning Routines for Effective Off-Policy Reinforcement LearningEdoardo Cetin, Oya ÇeliktutanICML 2021 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- Addressing Action Oscillations through Learning Policy InertiaChen Chen, Hongyao Tang, Jianye Hao, Wulong Liu 等AAAI 2021 · 被引用 27 次
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 被引用 150 次
- Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement LearningTianchen Zhou, Hairi, Haibo Yang, Jia Liu 等ICML 2024 · 被引用 4 次
- Finite-Time Convergence and Sample Complexity of Multi-Agent Actor-Critic Reinforcement Learning with Average RewardHairi, Jia Liu, Songtao LuICLR 2022 · 被引用 21 次
- Select before Act: Spatially Decoupled Action Repetition for Continuous ControlBuqing Nie, Yangqing Fu, Yue GaoICLR 2025
