Does "Do Differentiable Simulators Give Better Policy Gradients?" Give Better Policy Gradients?
Ku Onoda, Paavo Parmas, Manato Yaguchi, Yutaka Matsuo
摘要
In policy gradient reinforcement learning, access to a differentiable model enables 1st-order gradient estimation that accelerates learning compared to relying solely on derivative-free 0th-order estimators. However, discontinuous dynamics cause bias and undermine the effectiveness of 1st-order estimators. Prior work addressed this bias by constructing a confidence interval around the REINFORCE 0th-order gradient estimator and using these bounds to detect discontinuities. However, the REINFORCE estimator is notoriously noisy, and we find that this method requires task-specific hyperparameter tuning and has low sample efficiency. This paper asks whether such bias is the primary obstacle and what minimal fixes suffice. First, we re-examine standard discontinuous settings from prior work and introduce DDCG, a lightweight test that switches estimators in nonsmooth regions; with a single hyperparameter, DDCG achieves robust performance and remains reliable with small samples. Second, on differentiable robotics control tasks, we present IVW-H, a per-step inverse-variance implementation that stabilizes variance without explicit discontinuity detection and yields strong results. Together, these findings indicate that while estimator switching improves robustness in controlled studies, careful variance control often dominates in practical deployments.
Published as a conference paper at ICLR 2026 Now we make another assumption σ 2 ∆ < cV [f (x)], where c ∈ [0, 1]. Then we have the inequality
Gaussian distribution Eq. ( 34)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos 等ICLR 2022 · 被引用 141 次
- Do Differentiable Simulators Give Better Policy Gradients?Hyung Ju Terry Suh, Max Simchowitz, Kaiqing Zhang, Russ TedrakeICML 2022 · 被引用 129 次
- PODS: Policy Optimization via Differentiable SimulationMiguel Zamora, Momchil Peychev, Sehoon Ha, Martin T. Vechev 等ICML 2021 · 被引用 65 次
- Gradient Informed Proximal Policy OptimizationSanghyun Son, Laura Yu Zheng, Ryan Sullivan, Yi-Ling Qiao 等NeurIPS 2023 · 被引用 21 次
- Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact DynamicsShenao Zhang, Wanxin Jin, Zhaoran WangICML 2023 · 被引用 13 次
相关 Paper
- Stabilizing Policy Gradient Methods via Reward ProfilingShihab Ahmed, El Houcine Bergou, Yue Wang, Aritra DuttaAAAI 2026
- PAGE-PG: A Simple and Loopless Variance-Reduced Policy Gradient Method with Probabilistic Gradient EstimationMatilde Gargiani, Andrea Zanelli, Andrea Martinelli, Tyler H. Summers 等ICML 2022 · 被引用 17 次
- Time Discretization-Invariant Safe Action Repetition for Policy Gradient MethodsSeohong Park, Jaekyeom Kim, Gunhee KimNeurIPS 2021 · 被引用 33 次
- Adaptive Horizon Actor-Critic for Policy Learning in Contact-Rich Differentiable SimulationIgnat Georgiev, Krishnan Srinivasan, Jie Xu, Eric Heiden 等ICML 2024 · 被引用 27 次
- Distributions as Actions: A Unified Framework for Diverse Action SpacesJiamin He, A. Rupam Mahmood, Martha WhiteICLR 2026
