Addressing Action Oscillations through Learning Policy Inertia
Chen Chen, Hongyao Tang, Jianye Hao, Wulong Liu, Zhaopeng Meng
摘要
Deep reinforcement learning (DRL) algorithms have been demonstrated to be effective in a wide range of challenging decision making and control tasks. However, these methods typically suffer from severe action oscillations in particular in discrete action setting, which means that agents select different actions within consecutive steps even though states only slightly differ. This issue is often neglected since the policy is usually evaluated by its cumulative rewards only. Action oscillation strongly affects the user experience and can even cause serious potential security menace especially in realworld domains with the main concern of safety, such as autonomous driving. To this end, we introduce Policy Inertia Controller (PIC) which serves as a generic plug-in framework to off-the-shelf DRL algorithms, to enables adaptive trade-off between the optimality and smoothness of the learned policy in a formal way. We propose Nested Policy Iteration as a general training algorithm for PIC-augmented policy which ensures monotonically non-decreasing updates under some mild conditions. Further, we derive a practical DRL algorithm, namely Nested Soft Actor-Critic. Experiments on a collection of autonomous driving tasks and several Atari games suggest that our approach demonstrates substantial oscillation reduction in comparison to a range of commonly adopted baselines with almost no performance degradation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- TAAC: Temporally Abstract Actor-Critic for Continuous ControlHaonan Yu, Wei Xu, Haichao ZhangNeurIPS 2021 · 被引用 30 次
- LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal ControlXujie Song, Jingliang Duan, Wenxuan Wang, Shengbo Eben Li 等ICML 2023 · 被引用 19 次
- When to Sense and Control? A Time-adaptive Approach for Continuous-Time RLLenart Treven, Bhavya Sukhija, Yarden As, Florian Dörfler 等NeurIPS 2024 · 被引用 10 次
- Unlock the Intermittent Control Ability of Model Free Reinforcement LearningJiashun Liu, Jianye Hao, Xiaotian Hao, Yi Ma 等NeurIPS 2024 · 被引用 1 次
- Enhancing Control Policy Smoothness by Aligning Actions with Predictions from Preceding StatesKyoleen Kwak, Hyoseok HwangAAAI 2026 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- RVI-SAC: Average Reward Off-Policy Deep Reinforcement LearningYukinari Hisaki, Isao OnoICML 2024 · 被引用 6 次
- Stabilizing Policy Gradient Methods via Reward ProfilingShihab Ahmed, El Houcine Bergou, Yue Wang, Aritra DuttaAAAI 2026
- PiCor: Multi-Task Deep Reinforcement Learning with Policy CorrectionFengshuo Bai, Hongming Zhang, Tianyang Tao, Zhiheng Wu 等AAAI 2023 · 被引用 31 次
- Reinforcement Learning for Control with Multiple FrequenciesJongmin Lee, Byung-Jun Lee, Kee-Eung KimNeurIPS 2020 · 被引用 19 次
- Policy Gradient With Serial Markov Chain ReasoningEdoardo Cetin, Oya ÇeliktutanNeurIPS 2022 · 被引用 4 次
