Sampling Complexity of TD and PPO in RKHS
LU ZOU, Wendi Ren, WEIZHONG ZHANG, Liang Ding, Shuang Li
Abstract
We revisit Proximal Policy Optimization (PPO) from a function-space perspective. Our analysis decouples policy evaluation and improvement in a reproducing kernel Hilbert space (RKHS): (i) A kernelized temporal-difference (TD) critic performs efficient RKHS-gradient updates using only one-step state–action transition samples. (ii) a KL-regularized, natural-gradient policy step exponentiates the evaluated action-value, recovering a PPO/TRPO-style proximal update in continuous state-action spaces. We provide non-asymptotic, instance-adaptive guarantees whose rates depend on RKHS entropy, unifying tabular, linear, Sobolev, Gaussian, and Neural Tangent Kernel (NTK) regimes, and we derive a sampling rule for the proximal update that ensures the optimal convergence rate for stochastic optimization. Empirically, the theory-aligned schedule improves stability and sample efficiency on common control tasks (e.g., CartPole, Acrobot, and HalfCheetah), while our TD-based critic attains favorable throughput versus a GAE baseline. Altogether, our results place PPO on a firmer theoretical footing beyond finite-dimensional assumptions and clarify when RKHS-proximal updates with kernel-TD critics yield global policy improvement with practical efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 304 citations
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 201 citations
- Accountable Off-Policy Evaluation With Kernel Bellman StatisticsYihao Feng, Tongzheng Ren, Ziyang Tang, Qiang LiuICML 2020 · 45 citations
- Geometric Insights into the Convergence of Nonlinear TD LearningDavid Brandfonbrener, Joan BrunaICLR 2020 · 18 citations
Related papers
- A Non-asymptotic Analysis of Non-parametric Temporal-Difference LearningEloïse Berthier, Ziad Kobeissi, Francis R. BachNeurIPS 2022 · 6 citations
- Stabilizing PPO via Latent-Space Regularization and KDE-Driven ExplorationMeiyu Du, Yuqing Gao, Wei WangICML 2026
- Off-Policy Proximal Policy OptimizationWenjia Meng, Qian Zheng, Gang Pan, Yilong YinAAAI 2023 · 27 citations
- Reparameterization Proximal Policy OptimizationHai Zhong, Xun Wang, Zhuoran Li, Longbo HuangICML 2026
- PPO-Clip Attains Global Optimality: Towards Deeper Understandings of ClippingNai-Chieh Huang, Ping-Chun Hsieh, Kuo-Hao Ho, I-Chen WuAAAI 2024 · 34 citations
