Adaptive Horizon Actor-Critic for Policy Learning in Contact-Rich Differentiable Simulation
Ignat Georgiev, Krishnan Srinivasan, Jie Xu, Eric Heiden, Animesh Garg
Abstract
Model-Free Reinforcement Learning (MFRL), leveraging the policy gradient theorem, has demonstrated considerable success in continuous control tasks. However, these approaches are plagued by high gradient variance due to zerothorder gradient estimation, resulting in suboptimal policies. Conversely, First-Order Model-Based Reinforcement Learning (FO-MBRL) methods employing differentiable simulation provide gradients with reduced variance but are susceptible to sampling error in scenarios involving stiff dynamics, such as physical contact. This paper investigates the source of this error and introduces Adaptive Horizon Actor-Critic (AHAC), an FO-MBRL algorithm that reduces gradient error by adapting the model-based horizon to avoid stiff dynamics. Empirical findings reveal that AHAC outperforms MFRL baselines, attaining 40% more reward across a set of locomotion tasks and efficiently scaling to high-dimensional control environments with improved wall-clock-time efficiency. adaptive-horizon-actor-critic.github.io
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bb236c9-14eb-4188-845d-a6c7bdb68398Cited by top-tier papers8
- Differentiable Simulation of Hard Contacts with Soft Gradients for Learning and ControlAnselm Paulus, Andreas René Geist, Pierre Schumacher, Vít Musil et al.ICLR 2026 · 7 citations
- Relative Entropy Pathwise Policy OptimizationClaas Voelcker, Axel Brunnbauer, Marcel Hussing, Michal Nauman et al.ICLR 2026 · 6 citations
- Accelerated Learning with Linear Temporal Logic using Differentiable SimulationAlper Kamil Bozkurt, Calin Belta, Ming C. LinICLR 2026 · 2 citations
- Reparameterization Flow Policy OptimizationHai Zhong, Zhuoran Li, Xun Wang, Longbo HuangICML 2026
- Neural Control: Adjoint Learning Through Equilibrium ConstraintsDezhong Tong, Jiawen Wang, Hengyi Zhou, Yinlong Shen et al.ICML 2026
Builds on7
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos et al.ICLR 2022 · 141 citations
- Do Differentiable Simulators Give Better Policy Gradients?Hyung Ju Terry Suh, Max Simchowitz, Kaiqing Zhang, Russ TedrakeICML 2022 · 129 citations
Related papers
- Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact DynamicsShenao Zhang, Wanxin Jin, Zhaoran WangICML 2023 · 13 citations
- Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid ControlWeidong Huang, Zhehan Li, Hangxin Liu, Biao Hou et al.ICLR 2026 · 4 citations
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 164 citations
- Adaptive Reinforcement Learning for Unobservable Random DelaysJohn Wikman, Alexandre Proutiere, David BromanICML 2026
- PWM: Policy Learning with Multi-Task World ModelsIgnat Georgiev, Varun Giridhar, Nicklas Hansen, Animesh GargICLR 2025
