Residual Kernel Policy Network: Enhancing Stability and Robustness in RKHS-Based Reinforcement Learning
Yixian Zhang, Huaze Tang, Huijing Lin, Wenbo Ding
摘要
Achieving optimal performance in reinforcement learning requires robust policies supported by training processes that ensure both sample efficiency and stability. Modeling the policy in reproducing kernel Hilbert space (RKHS) enables efficient exploration of local optimal solutions. However, the stability of existing RKHSbased methods is hindered by significant variance in gradients, while the robustness of the learned policies is often compromised due to the sensitivity of hyperparameters. In this work, we conduct a comprehensive analysis of the significant instability in RKHS policies and reveal that the variance of the policy gradient increases substantially when a wide-bandwidth kernel is employed. To address these challenges, we propose a novel RKHS policy learning method integrated with representation learning to dynamically process observations in complex environments, enhancing the robustness of RKHS policies. Furthermore, inspired by the advantage functions, we introduce a residual layer that further stabilizes the training process by significantly reducing gradient variance in RKHS. Our novel algorithm, the Residual Kernel Policy Network (ResKPN), demonstrates state-of-the-art performance, achieving a 30% improvement in episodic rewards across complex environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU NetworksAgustinus Kristiadi, Matthias Hein, Philipp HennigICML 2020 · 被引用 344 次
- Discovered Policy OptimisationChris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz 等NeurIPS 2022 · 被引用 134 次
- Coordinated Proximal Policy OptimizationZifan Wu, Chao Yu, Deheng Ye, Junge Zhang 等NeurIPS 2021 · 被引用 73 次
- Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement LearningSumeet Batra, Bryon Tjanaka, Matthew Christopher Fontaine, Aleksei Petrenko 等ICLR 2024 · 被引用 26 次
- Cliff Diving: Exploring Reward Surfaces in Reinforcement Learning EnvironmentsRyan Sullivan, J. K. Terry, Benjamin Black, John P. DickersonICML 2022 · 被引用 11 次
相关 Paper
- Policy Newton Algorithm in Reproducing Kernel Hilbert SpaceYixian Zhang, Huaze Tang, Changxu Wei, Chao Wang 等ICLR 2026 · 被引用 3 次
- Kernelized Reinforcement Learning with Order Optimal Regret BoundsSattar Vakili, Julia OlkhovskayaNeurIPS 2023 · 被引用 22 次
- Sampling Complexity of TD and PPO in RKHSLU ZOU, Wendi Ren, WEIZHONG ZHANG, Liang Ding 等ICLR 2026 · 被引用 1 次
- SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement LearningNa Li, Hangguan Shan, Wei Ni, Wenjie Zhang 等ICML 2026
- A Non-asymptotic Analysis of Non-parametric Temporal-Difference LearningEloïse Berthier, Ziad Kobeissi, Francis R. BachNeurIPS 2022 · 被引用 6 次
