Accelerated Policy Learning with Parallel Differentiable Simulation
Jie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos, Wojciech Matusik, Animesh Garg, Miles Macklin
摘要
Deep reinforcement learning can generate complex control policies, but requires large amounts of training data to work effectively. Recent work has attempted to address this issue by leveraging differentiable simulators. However, inherent problems such as local minima and exploding/vanishing numerical gradients prevent these methods from being generally applied to control tasks with complex contact-rich dynamics, such as humanoid locomotion in classical RL benchmarks. In this work we present a high-performance differentiable simulator and a new policy learning algorithm (SHAC) that can effectively leverage simulation gradients, even in the presence of non-smoothness. Our learning algorithm alleviates problems with local minima through a smooth critic function, avoids vanishing/exploding gradients through a truncated learning window, and allows many physical environments to be run in parallel. We evaluate our method on classical RL control tasks, and show substantial improvements in sample efficiency and wall-clock time over state-of-the-art RL and differentiable simulation-based algorithms. In addition, we demonstrate the scalability of our method by applying it to the challenging high-dimensional problem of muscle-actuated locomotion with a large action space, achieving a greater than 17× reduction in training time over the best-performing established RL algorithm. More visual results are provided at: https://short-horizon-actor-critic.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper48
- Does "Do Differentiable Simulators Give Better Policy Gradients?" Give Better Policy Gradients?Ku Onoda, Paavo Parmas, Manato Yaguchi, Yutaka MatsuoICLR 2026 · 被引用 134 次
- NeuPhysics: Editable Neural Geometry and Physics from Monocular VideosYi-Ling Qiao, Alexander Gao, Ming C. LinNeurIPS 2022 · 被引用 62 次
- PPR: Physically Plausible Reconstruction from Monocular VideosGengshan Yang, Shuo Yang, John Z. Zhang, Zachary Manchester 等ICCV 2023 · 被引用 41 次
- Adaptive Horizon Actor-Critic for Policy Learning in Contact-Rich Differentiable SimulationIgnat Georgiev, Krishnan Srinivasan, Jie Xu, Eric Heiden 等ICML 2024 · 被引用 27 次
- Gradient Informed Proximal Policy OptimizationSanghyun Son, Laura Yu Zheng, Ryan Sullivan, Yi-Ling Qiao 等NeurIPS 2023 · 被引用 21 次
它引用的顶会 Paper8
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- DiffTaichi: Differentiable Programming for Physical SimulationYuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun 等ICLR 2020 · 被引用 479 次
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine 等SIGGRAPH 2021 · 被引用 392 次
- PlasticineLab: A Soft-Body Manipulation Benchmark with Differentiable PhysicsZhiao Huang, Yuanming Hu, Tao Du, Siyuan Zhou 等ICLR 2021 · 被引用 164 次
- Model-Augmented Actor-Critic: Backpropagating through PathsIgnasi Clavera, Yao Fu, Pieter AbbeelICLR 2020 · 被引用 96 次
相关 Paper
- Accelerating Visual-Policy Learning through Parallel Differentiable SimulationHaoxiang You, Yilang Liu, Ian AbrahamNeurIPS 2025 · 被引用 7 次
- Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid ControlWeidong Huang, Zhehan Li, Hangxin Liu, Biao Hou 等ICLR 2026 · 被引用 4 次
- Do Differentiable Simulators Give Better Policy Gradients?Hyung Ju Terry Suh, Max Simchowitz, Kaiqing Zhang, Russ TedrakeICML 2022 · 被引用 129 次
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 被引用 164 次
- Efficient Differentiable Simulation of Articulated BodiesYi-Ling Qiao, Junbang Liang, Vladlen Koltun, Ming C. LinICML 2021 · 被引用 68 次
