Residual Kernel Policy Network: Enhancing Stability and Robustness in RKHS-Based Reinforcement Learning
Yixian Zhang, Huaze Tang, Huijing Lin, Wenbo Ding
Abstract
Achieving optimal performance in reinforcement learning requires robust policies supported by training processes that ensure both sample efficiency and stability. Modeling the policy in reproducing kernel Hilbert space (RKHS) enables efficient exploration of local optimal solutions. However, the stability of existing RKHSbased methods is hindered by significant variance in gradients, while the robustness of the learned policies is often compromised due to the sensitivity of hyperparameters. In this work, we conduct a comprehensive analysis of the significant instability in RKHS policies and reveal that the variance of the policy gradient increases substantially when a wide-bandwidth kernel is employed. To address these challenges, we propose a novel RKHS policy learning method integrated with representation learning to dynamically process observations in complex environments, enhancing the robustness of RKHS policies. Furthermore, inspired by the advantage functions, we introduce a residual layer that further stabilizes the training process by significantly reducing gradient variance in RKHS. Our novel algorithm, the Residual Kernel Policy Network (ResKPN), demonstrates state-of-the-art performance, achieving a 30% improvement in episodic rewards across complex environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU NetworksAgustinus Kristiadi, Matthias Hein, Philipp HennigICML 2020 · 344 citations
- Discovered Policy OptimisationChris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz et al.NeurIPS 2022 · 134 citations
- Coordinated Proximal Policy OptimizationZifan Wu, Chao Yu, Deheng Ye, Junge Zhang et al.NeurIPS 2021 · 73 citations
- Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement LearningSumeet Batra, Bryon Tjanaka, Matthew Christopher Fontaine, Aleksei Petrenko et al.ICLR 2024 · 26 citations
- Cliff Diving: Exploring Reward Surfaces in Reinforcement Learning EnvironmentsRyan Sullivan, J. K. Terry, Benjamin Black, John P. DickersonICML 2022 · 11 citations
Related papers
- Policy Newton Algorithm in Reproducing Kernel Hilbert SpaceYixian Zhang, Huaze Tang, Changxu Wei, Chao Wang et al.ICLR 2026 · 3 citations
- Kernelized Reinforcement Learning with Order Optimal Regret BoundsSattar Vakili, Julia OlkhovskayaNeurIPS 2023 · 22 citations
- Sampling Complexity of TD and PPO in RKHSLU ZOU, Wendi Ren, WEIZHONG ZHANG, Liang Ding et al.ICLR 2026 · 1 citation
- SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement LearningNa Li, Hangguan Shan, Wei Ni, Wenjie Zhang et al.ICML 2026
- A Non-asymptotic Analysis of Non-parametric Temporal-Difference LearningEloïse Berthier, Ziad Kobeissi, Francis R. BachNeurIPS 2022 · 6 citations
