Sampling Complexity of TD and PPO in RKHS
LU ZOU, Wendi Ren, WEIZHONG ZHANG, Liang Ding, Shuang Li
摘要
We revisit Proximal Policy Optimization (PPO) from a function-space perspective. Our analysis decouples policy evaluation and improvement in a reproducing kernel Hilbert space (RKHS): (i) A kernelized temporal-difference (TD) critic performs efficient RKHS-gradient updates using only one-step state–action transition samples. (ii) a KL-regularized, natural-gradient policy step exponentiates the evaluated action-value, recovering a PPO/TRPO-style proximal update in continuous state-action spaces. We provide non-asymptotic, instance-adaptive guarantees whose rates depend on RKHS entropy, unifying tabular, linear, Sobolev, Gaussian, and Neural Tangent Kernel (NTK) regimes, and we derive a sampling rule for the proximal update that ensures the optimal convergence rate for stochastic optimization. Empirically, the theory-aligned schedule improves stability and sample efficiency on common control tasks (e.g., CartPole, Acrobot, and HalfCheetah), while our TD-based critic attains favorable throughput versus a GAE baseline. Altogether, our results place PPO on a firmer theoretical footing beyond finite-dimensional assumptions and clarify when RKHS-proximal updates with kernel-TD critics yield global policy improvement with practical efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 被引用 304 次
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 被引用 201 次
- Accountable Off-Policy Evaluation With Kernel Bellman StatisticsYihao Feng, Tongzheng Ren, Ziyang Tang, Qiang LiuICML 2020 · 被引用 45 次
- Geometric Insights into the Convergence of Nonlinear TD LearningDavid Brandfonbrener, Joan BrunaICLR 2020 · 被引用 18 次
相关 Paper
- A Non-asymptotic Analysis of Non-parametric Temporal-Difference LearningEloïse Berthier, Ziad Kobeissi, Francis R. BachNeurIPS 2022 · 被引用 6 次
- Stabilizing PPO via Latent-Space Regularization and KDE-Driven ExplorationMeiyu Du, Yuqing Gao, Wei WangICML 2026
- Off-Policy Proximal Policy OptimizationWenjia Meng, Qian Zheng, Gang Pan, Yilong YinAAAI 2023 · 被引用 27 次
- Reparameterization Proximal Policy OptimizationHai Zhong, Xun Wang, Zhuoran Li, Longbo HuangICML 2026
- PPO-Clip Attains Global Optimality: Towards Deeper Understandings of ClippingNai-Chieh Huang, Ping-Chun Hsieh, Kuo-Hao Ho, I-Chen WuAAAI 2024 · 被引用 34 次
