PODS: Policy Optimization via Differentiable Simulation
Miguel Zamora, Momchil Peychev, Sehoon Ha, Martin T. Vechev, Stelian Coros
摘要
Current reinforcement learning (RL) methods use simulation models as simple black-box oracles. In this paper, with the goal of improving the performance exhibited by RL algorithms, we explore a systematic way of leveraging the additional information provided by an emerging class of differentiable simulators. Building on concepts established by Deterministic Policy Gradients (DPG) methods, the neural network policies learned with our approach represent deterministic actions. In a departure from standard methodologies, however, learning these policies does not hinge on approximations of the value function that must be learned concurrently in an actor-critic fashion. Instead, we exploit differentiable simulators to directly compute the analytic gradient of a policy's value function with respect to the actions it outputs. This, in turn, allows us to effciently perform locally optimal policy improvement iterations. Compared against other state-of-the-art RL methods, we show that with minimal hyperparameter tuning our approach consistently leads to better asymptotic behavior across a set of payload manipulation tasks that demand a high degree of accuracy and precision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos 等ICLR 2022 · 被引用 141 次
- Does "Do Differentiable Simulators Give Better Policy Gradients?" Give Better Policy Gradients?Ku Onoda, Paavo Parmas, Manato Yaguchi, Yutaka MatsuoICLR 2026 · 被引用 134 次
- Gradient Informed Proximal Policy OptimizationSanghyun Son, Laura Yu Zheng, Ryan Sullivan, Yi-Ling Qiao 等NeurIPS 2023 · 被引用 21 次
- FluidLab: A Differentiable Environment for Benchmarking Complex Fluid ManipulationZhou Xian, Bo Zhu, Zhenjia Xu, Hsiao-Yu Tung 等ICLR 2023 · 被引用 11 次
- Thin-Shell Object Manipulations With Differentiable Physics SimulationsYian Wang, Juntian Zheng, Zhehuan Chen, Zhou Xian 等ICLR 2024 · 被引用 10 次
它引用的顶会 Paper2
相关 Paper
- Unlocking Efficient Vehicle Dynamics Modeling via Analytic World ModelsAsen Nachkov, Danda Pani Paudel, Jan-Nico Zaech, Davide Scaramuzza 等AAAI 2026 · 被引用 2 次
- Deterministic Value-Policy GradientsQingpeng Cai, Ling Pan, Pingzhong TangAAAI 2020 · 被引用 1 次
- Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable SimulationsFeng Gao, Liangzhi Shi, Shenao Zhang, Zhaoran Wang 等ICML 2024 · 被引用 7 次
- DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit DifferentiationShuyuan Wang, Philip D. Loewen, Michael G. Forbes, R. Bhushan Gopaluni 等ICML 2025
- DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with ToolsXingyu Lin, Zhiao Huang, Yunzhu Li, Joshua B. Tenenbaum 等ICLR 2022 · 被引用 85 次
