Do Differentiable Simulators Give Better Policy Gradients?
Hyung Ju Terry Suh, Max Simchowitz, Kaiqing Zhang, Russ Tedrake
Abstract
Differentiable simulators promise faster computation time for reinforcement learning by replacing zeroth-order gradient estimates of a stochastic objective with an estimate based on first-order gradients. However, it is yet unclear what factors decide the performance of the two estimators on complex landscapes that involve long-horizon planning and control on physical systems, despite the crucial relevance of this question for the utility of differentiable simulators. We show that characteristics of certain physical systems, such as stiffness or discontinuities, may compromise the efficacy of the first-order estimator, and analyze this phenomenon through the lens of bias and variance. We additionally propose an -order gradient estimator, with , which correctly utilizes exact gradients to combine the efficiency of first-order estimates with the robustness of zero-order methods. We demonstrate the pitfalls of traditional estimators and the advantages of the -order estimator on some numerical examples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e79589f4-9562-4490-871e-d9ae709d6077Cited by top-tier papers43
- Does "Do Differentiable Simulators Give Better Policy Gradients?" Give Better Policy Gradients?Ku Onoda, Paavo Parmas, Manato Yaguchi, Yutaka MatsuoICLR 2026 · 134 citations
- Adaptive Horizon Actor-Critic for Policy Learning in Contact-Rich Differentiable SimulationIgnat Georgiev, Krishnan Srinivasan, Jie Xu, Eric Heiden et al.ICML 2024 · 27 citations
- Gradient Informed Proximal Policy OptimizationSanghyun Son, Laura Yu Zheng, Ryan Sullivan, Yi-Ling Qiao et al.NeurIPS 2023 · 21 citations
- Differentiable Simulations for Enhanced Sampling of Rare EventsMartin Sípka, Johannes C. B. Dietschreit, Lukás Grajciar, Rafael Gómez-BombarelliICML 2023 · 17 citations
- Smoothed Online Learning for Prediction in Piecewise Affine SystemsAdam Block, Max Simchowitz, Russ TedrakeNeurIPS 2023 · 13 citations
Builds on2
- PlasticineLab: A Soft-Body Manipulation Benchmark with Differentiable PhysicsZhiao Huang, Yuanming Hu, Tao Du, Siyuan Zhou et al.ICLR 2021 · 164 citations
- Systematically differentiating parametric discontinuitiesSai Praveen Bangaru, Jesse Michel, Kevin Mu, Gilbert Bernstein et al.SIGGRAPH 2021 · 30 citations
Related papers
- Enabling First-Order Gradient-Based Learning for Equilibrium Computation in MarketsNils Kohring, Fabian Raoul Pieroth, Martin BichlerICML 2023 · 8 citations
- Efficient Differentiable Contact Model with Long-range InfluenceXiaohan Ye, Kui Wu, Taku Komura, Zherong PanICLR 2026 · 2 citations
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos et al.ICLR 2022 · 141 citations
- DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with ToolsXingyu Lin, Zhiao Huang, Yunzhu Li, Joshua B. Tenenbaum et al.ICLR 2022 · 85 citations
- Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact DynamicsShenao Zhang, Wanxin Jin, Zhaoran WangICML 2023 · 13 citations
