Enhancing Control Policy Smoothness by Aligning Actions with Predictions from Preceding States
Kyoleen Kwak, Hyoseok Hwang
Abstract
Deep reinforcement learning has proven to be a powerful approach to solving control tasks, but its characteristic high‑frequency oscillations make it difficult to apply in real‑world environments. While prior methods have addressed action oscillations via architectural or loss-based methods, the latter typically depend on heuristic or synthetic definitions of state similarity to promote action consistency, which often fail to accurately reflect the underlying system dynamics. In this paper, we propose a novel loss-based method by introducing a transition-induced similar state. The transition-induced similar state is defined as the distribution of next states transitioned from the previous state. Since it utilizes only environmental feedback and actually collected data, it better captures system dynamics. Building upon this foundation, we introduce Action Smoothing by Aligning Actions with Predictions from Preceding States (ASAP), an action smoothing method that effectively mitigates action oscillations. ASAP enforces action smoothness by aligning the actions with those taken in transition-induced similar states and by penalizing second-order differences to suppress high-frequency oscillations. Experiments in Gymnasium and Isaac-lab environments demonstrate that ASAP yields smoother control and improved policy performance over existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92e498e6-949e-42a7-9b35-a4afdc095b83Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Spectral Normalisation for Deep Reinforcement Learning: An Optimisation PerspectiveFlorin Gogianu, Tudor Berariu, Mihaela Rosca, Claudia Clopath et al.ICML 2021 · 69 citations
- Addressing Action Oscillations through Learning Policy InertiaChen Chen, Hongyao Tang, Jianye Hao, Wulong Liu et al.AAAI 2021 · 27 citations
- LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal ControlXujie Song, Jingliang Duan, Wenxuan Wang, Shengbo Eben Li et al.ICML 2023 · 19 citations
- LipsNet++: Unifying Filter and Controller into a Policy NetworkXujie Song, Liangfa Chen, Tong Liu, Wenxuan Wang et al.ICML 2025
Related papers
- Implicit Action Chunking for Smooth Continuous ControlBosun Liang, Shuo Pei, Zirui Chen, Chuanzhi Fan et al.ICML 2026
- Deep Reinforcement Learning with Robust and Smooth PolicyQianli Shen, Yan Li, Haoming Jiang, Zhaoran Wang et al.ICML 2020 · 95 citations
- Learning Generalizable Representations for Reinforcement Learning via Adaptive Meta-learner of Behavioral SimilaritiesJianda Chen, Sinno Jialin PanICLR 2022 · 6 citations
- ODE-based Smoothing Neural Network for Reinforcement Learning TasksYinuo Wang, Wenxuan Wang, Xujie Song, Tong Liu et al.ICLR 2025
- Model-based Reinforcement Learning for Parameterized Action SpacesRenhao Zhang, Haotian Fu, Yilin Miao, George KonidarisICML 2024 · 8 citations
