Addressing Action Oscillations through Learning Policy Inertia
Chen Chen, Hongyao Tang, Jianye Hao, Wulong Liu, Zhaopeng Meng
Abstract
Deep reinforcement learning (DRL) algorithms have been demonstrated to be effective in a wide range of challenging decision making and control tasks. However, these methods typically suffer from severe action oscillations in particular in discrete action setting, which means that agents select different actions within consecutive steps even though states only slightly differ. This issue is often neglected since the policy is usually evaluated by its cumulative rewards only. Action oscillation strongly affects the user experience and can even cause serious potential security menace especially in realworld domains with the main concern of safety, such as autonomous driving. To this end, we introduce Policy Inertia Controller (PIC) which serves as a generic plug-in framework to off-the-shelf DRL algorithms, to enables adaptive trade-off between the optimality and smoothness of the learned policy in a formal way. We propose Nested Policy Iteration as a general training algorithm for PIC-augmented policy which ensures monotonically non-decreasing updates under some mild conditions. Further, we derive a practical DRL algorithm, namely Nested Soft Actor-Critic. Experiments on a collection of autonomous driving tasks and several Atari games suggest that our approach demonstrates substantial oscillation reduction in comparison to a range of commonly adopted baselines with almost no performance degradation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext deb45155-a259-440c-a8fe-c299bcf6903bCited by top-tier papers9
- TAAC: Temporally Abstract Actor-Critic for Continuous ControlHaonan Yu, Wei Xu, Haichao ZhangNeurIPS 2021 · 30 citations
- LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal ControlXujie Song, Jingliang Duan, Wenxuan Wang, Shengbo Eben Li et al.ICML 2023 · 19 citations
- When to Sense and Control? A Time-adaptive Approach for Continuous-Time RLLenart Treven, Bhavya Sukhija, Yarden As, Florian Dörfler et al.NeurIPS 2024 · 10 citations
- Unlock the Intermittent Control Ability of Model Free Reinforcement LearningJiashun Liu, Jianye Hao, Xiaotian Hao, Yi Ma et al.NeurIPS 2024 · 1 citation
- Enhancing Control Policy Smoothness by Aligning Actions with Predictions from Preceding StatesKyoleen Kwak, Hyoseok HwangAAAI 2026 · 1 citation
Builds on1
Related papers
- RVI-SAC: Average Reward Off-Policy Deep Reinforcement LearningYukinari Hisaki, Isao OnoICML 2024 · 6 citations
- Stabilizing Policy Gradient Methods via Reward ProfilingShihab Ahmed, El Houcine Bergou, Yue Wang, Aritra DuttaAAAI 2026
- PiCor: Multi-Task Deep Reinforcement Learning with Policy CorrectionFengshuo Bai, Hongming Zhang, Tianyang Tao, Zhiheng Wu et al.AAAI 2023 · 31 citations
- Reinforcement Learning for Control with Multiple FrequenciesJongmin Lee, Byung-Jun Lee, Kee-Eung KimNeurIPS 2020 · 19 citations
- Policy Gradient With Serial Markov Chain ReasoningEdoardo Cetin, Oya ÇeliktutanNeurIPS 2022 · 4 citations
