Gradient Informed Proximal Policy Optimization
Sanghyun Son, Laura Yu Zheng, Ryan Sullivan, Yi-Ling Qiao, Ming C. Lin
摘要
We introduce a novel policy learning method that integrates analytical gradients from differentiable environments with the Proximal Policy Optimization (PPO) algorithm. To incorporate analytical gradients into the PPO framework, we introduce the concept of an α-policy that stands as a locally superior policy. By adaptively modifying the α value, we can effectively manage the influence of analytical policy gradients during learning. To this end, we suggest metrics for assessing the variance and bias of analytical gradients, reducing dependence on these gradients when high variance or bias is detected. Our proposed approach outperforms baseline algorithms in various scenarios, such as function optimization, physics simulations, and traffic control environments. Our code can be found online: https://github.com/SonSang/gippo . 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Does "Do Differentiable Simulators Give Better Policy Gradients?" Give Better Policy Gradients?Ku Onoda, Paavo Parmas, Manato Yaguchi, Yutaka MatsuoICLR 2026 · 被引用 134 次
- Relative Entropy Pathwise Policy OptimizationClaas Voelcker, Axel Brunnbauer, Marcel Hussing, Michal Nauman 等ICLR 2026 · 被引用 6 次
- Accelerated Learning with Linear Temporal Logic using Differentiable SimulationAlper Kamil Bozkurt, Calin Belta, Ming C. LinICLR 2026 · 被引用 2 次
- Reparameterization Proximal Policy OptimizationHai Zhong, Xun Wang, Zhuoran Li, Longbo HuangICML 2026
- Stabilizing Reinforcement Learning in Differentiable Multiphysics SimulationEliot Xing, Vernon Luk, Jean OhICLR 2025
它引用的顶会 Paper4
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos 等ICLR 2022 · 被引用 141 次
- Do Differentiable Simulators Give Better Policy Gradients?Hyung Ju Terry Suh, Max Simchowitz, Kaiqing Zhang, Russ TedrakeICML 2022 · 被引用 129 次
- Efficient Differentiable Simulation of Articulated BodiesYi-Ling Qiao, Junbang Liang, Vladlen Koltun, Ming C. LinICML 2021 · 被引用 68 次
- PODS: Policy Optimization via Differentiable SimulationMiguel Zamora, Momchil Peychev, Sehoon Ha, Martin T. Vechev 等ICML 2021 · 被引用 65 次
相关 Paper
- Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable SimulationsFeng Gao, Liangzhi Shi, Shenao Zhang, Zhaoran Wang 等ICML 2024 · 被引用 7 次
- Gradient Information Matters in Policy Optimization by Back-propagating through ModelChongchong Li, Yue Wang, Wei Chen, Yuting Liu 等ICLR 2022 · 被引用 10 次
- On Proximal Policy Optimization's Heavy-tailed GradientsSaurabh Garg, Joshua Zhanson, Emilio Parisotto, Adarsh Prasad 等ICML 2021 · 被引用 32 次
- Off-Policy Proximal Policy OptimizationWenjia Meng, Qian Zheng, Gang Pan, Yilong YinAAAI 2023 · 被引用 27 次
- What about Inputting Policy in Value Function: Policy Representation and Policy-Extended Value Function ApproximatorHongyao Tang, Zhaopeng Meng, Jianye Hao, Chen Chen 等AAAI 2022 · 被引用 20 次
