PODS: Policy Optimization via Differentiable Simulation
Miguel Zamora, Momchil Peychev, Sehoon Ha, Martin T. Vechev, Stelian Coros
Abstract
Current reinforcement learning (RL) methods use simulation models as simple black-box oracles. In this paper, with the goal of improving the performance exhibited by RL algorithms, we explore a systematic way of leveraging the additional information provided by an emerging class of differentiable simulators. Building on concepts established by Deterministic Policy Gradients (DPG) methods, the neural network policies learned with our approach represent deterministic actions. In a departure from standard methodologies, however, learning these policies does not hinge on approximations of the value function that must be learned concurrently in an actor-critic fashion. Instead, we exploit differentiable simulators to directly compute the analytic gradient of a policy's value function with respect to the actions it outputs. This, in turn, allows us to effciently perform locally optimal policy improvement iterations. Compared against other state-of-the-art RL methods, we show that with minimal hyperparameter tuning our approach consistently leads to better asymptotic behavior across a set of payload manipulation tasks that demand a high degree of accuracy and precision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos et al.ICLR 2022 · 141 citations
- Does "Do Differentiable Simulators Give Better Policy Gradients?" Give Better Policy Gradients?Ku Onoda, Paavo Parmas, Manato Yaguchi, Yutaka MatsuoICLR 2026 · 134 citations
- Gradient Informed Proximal Policy OptimizationSanghyun Son, Laura Yu Zheng, Ryan Sullivan, Yi-Ling Qiao et al.NeurIPS 2023 · 21 citations
- FluidLab: A Differentiable Environment for Benchmarking Complex Fluid ManipulationZhou Xian, Bo Zhu, Zhenjia Xu, Hsiao-Yu Tung et al.ICLR 2023 · 11 citations
- Thin-Shell Object Manipulations With Differentiable Physics SimulationsYian Wang, Juntian Zheng, Zhehuan Chen, Zhou Xian et al.ICLR 2024 · 10 citations
Builds on2
Related papers
- Unlocking Efficient Vehicle Dynamics Modeling via Analytic World ModelsAsen Nachkov, Danda Pani Paudel, Jan-Nico Zaech, Davide Scaramuzza et al.AAAI 2026 · 2 citations
- Deterministic Value-Policy GradientsQingpeng Cai, Ling Pan, Pingzhong TangAAAI 2020 · 1 citation
- Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable SimulationsFeng Gao, Liangzhi Shi, Shenao Zhang, Zhaoran Wang et al.ICML 2024 · 7 citations
- DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit DifferentiationShuyuan Wang, Philip D. Loewen, Michael G. Forbes, R. Bhushan Gopaluni et al.ICML 2025
- DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with ToolsXingyu Lin, Zhiao Huang, Yunzhu Li, Joshua B. Tenenbaum et al.ICLR 2022 · 85 citations
