Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact Dynamics
Shenao Zhang, Wanxin Jin, Zhaoran Wang
摘要
Differentiable physics-based simulators have witnessed remarkable success in robot learning involving contact dynamics, benefiting from their improved accuracy and efficiency in solving the underlying complementarity problem. However, when utilizing the First-Order Policy Gradient (FOPG) method, our theory indicates that the complementarity-based systems suffer from stiffness, leading to an explosion in the gradient variance of FOPG. As a result, optimization becomes challenging due to chaotic and non-smooth loss landscapes. To tackle this issue, we propose a novel approach called Adaptive Barrier Smoothing (ABS), which introduces a class of softened complementarity systems that correspond to barrier-smoothed objectives. With a contact-aware adaptive central-path parameter, ABS reduces the FOPG gradient variance while controlling the gradient bias. We justify the adaptive design by analyzing the roots of the system's stiffness. Additionally, we establish the convergence of FOPG and show that ABS achieves a reasonable trade-off between the gradient variance and bias by providing their upper bounds. Moreover, we present a variant of FOPG based on complementarity modeling that efficiently fits the contact dynamics by learning the physical parameters. Experimental results on various robotic tasks are provided to support our theory and method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Does "Do Differentiable Simulators Give Better Policy Gradients?" Give Better Policy Gradients?Ku Onoda, Paavo Parmas, Manato Yaguchi, Yutaka MatsuoICLR 2026 · 被引用 134 次
- Model-Based Reparameterization Policy Gradient Methods: Theory and Practical AlgorithmsShenao Zhang, Boyi Liu, Zhaoran Wang, Tuo ZhaoNeurIPS 2023 · 被引用 8 次
- Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable SimulationsFeng Gao, Liangzhi Shi, Shenao Zhang, Zhaoran Wang 等ICML 2024 · 被引用 7 次
- Mollification Effects of Policy Gradient MethodsTao Wang, Sylvia L. Herbert, Sicun GaoICML 2024 · 被引用 2 次
- Reparameterization Proximal Policy OptimizationHai Zhong, Xun Wang, Zhuoran Li, Longbo HuangICML 2026
它引用的顶会 Paper9
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 被引用 270 次
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 被引用 164 次
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos 等ICLR 2022 · 被引用 141 次
- Do Differentiable Simulators Give Better Policy Gradients?Hyung Ju Terry Suh, Max Simchowitz, Kaiqing Zhang, Russ TedrakeICML 2022 · 被引用 129 次
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 被引用 120 次
相关 Paper
- Adaptive Horizon Actor-Critic for Policy Learning in Contact-Rich Differentiable SimulationIgnat Georgiev, Krishnan Srinivasan, Jie Xu, Eric Heiden 等ICML 2024 · 被引用 27 次
- Efficient Differentiable Contact Model with Long-range InfluenceXiaohan Ye, Kui Wu, Taku Komura, Zherong PanICLR 2026 · 被引用 2 次
- Differentiable Simulation of Hard Contacts with Soft Gradients for Learning and ControlAnselm Paulus, Andreas René Geist, Pierre Schumacher, Vít Musil 等ICLR 2026 · 被引用 7 次
- Augmented Bayesian Policy SearchMahdi Kallel, Debabrota Basu, Riad Akrour, Carlo D'EramoICLR 2024 · 被引用 4 次
- Primal-Dual Non-Smooth Friction for Rigid Body AnimationYi-Lu Chen, Mickaël Ly, Chris WojtanSIGGRAPH 2024 · 被引用 4 次
