Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable Simulations
Feng Gao, Liangzhi Shi, Shenao Zhang, Zhaoran Wang, Yi Wu
摘要
Recent advancements in differentiable simulators highlight the potential of policy optimization using simulation gradients. Yet, these approaches are largely contingent on the continuity and smoothness of the simulation, which precludes the use of certain simulation engines, such as Mujoco. To tackle this challenge, we introduce the adaptive analytic gradient. This method views the Q function as a surrogate for future returns, consistent with the Bellman equation. By analyzing the variance of batched gradients, our method can autonomously opt for a more resilient Q function to compute the gradient when encountering rough simulation transitions. We also put forth the Adaptive-Gradient Policy Optimization (AGPO) algorithm, which leverages our proposed method for policy learning. On the theoretical side, we demonstrate AGPO's convergence, emphasizing its stable performance under non-smooth dynamics due to low variance. On the empirical side, our results show that AGPO effectively mitigates the challenges posed by non-smoothness in policy learning through differentiable simulation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Does "Do Differentiable Simulators Give Better Policy Gradients?" Give Better Policy Gradients?Ku Onoda, Paavo Parmas, Manato Yaguchi, Yutaka MatsuoICLR 2026 · 被引用 134 次
- Reparameterization Flow Policy OptimizationHai Zhong, Zhuoran Li, Xun Wang, Longbo HuangICML 2026
- Neural Control: Adjoint Learning Through Equilibrium ConstraintsDezhong Tong, Jiawen Wang, Hengyi Zhou, Yinlong Shen 等ICML 2026
- Reparameterization Proximal Policy OptimizationHai Zhong, Xun Wang, Zhuoran Li, Longbo HuangICML 2026
- Stabilizing Reinforcement Learning in Differentiable Multiphysics SimulationEliot Xing, Vernon Luk, Jean OhICLR 2025
它引用的顶会 Paper9
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine 等SIGGRAPH 2021 · 被引用 392 次
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos 等ICLR 2022 · 被引用 141 次
- Do Differentiable Simulators Give Better Policy Gradients?Hyung Ju Terry Suh, Max Simchowitz, Kaiqing Zhang, Russ TedrakeICML 2022 · 被引用 129 次
- Efficient Differentiable Simulation of Articulated BodiesYi-Ling Qiao, Junbang Liang, Vladlen Koltun, Ming C. LinICML 2021 · 被引用 68 次
- Learning Physically Simulated Tennis Skills from Broadcast VideosHaotian Zhang, Ye Yuan, Viktor Makoviychuk, Yunrong Guo 等SIGGRAPH 2023 · 被引用 47 次
相关 Paper
- Gradient Informed Proximal Policy OptimizationSanghyun Son, Laura Yu Zheng, Ryan Sullivan, Yi-Ling Qiao 等NeurIPS 2023 · 被引用 21 次
- Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact DynamicsShenao Zhang, Wanxin Jin, Zhaoran WangICML 2023 · 被引用 13 次
- PODS: Policy Optimization via Differentiable SimulationMiguel Zamora, Momchil Peychev, Sehoon Ha, Martin T. Vechev 等ICML 2021 · 被引用 65 次
- Unlocking Efficient Vehicle Dynamics Modeling via Analytic World ModelsAsen Nachkov, Danda Pani Paudel, Jan-Nico Zaech, Davide Scaramuzza 等AAAI 2026 · 被引用 2 次
- Enabling First-Order Gradient-Based Learning for Equilibrium Computation in MarketsNils Kohring, Fabian Raoul Pieroth, Martin BichlerICML 2023 · 被引用 8 次
