Extracting Strong Policies for Robotics Tasks from Zero-Order Trajectory Optimizers
Cristina Pinneri, Shambhuraj Sawant, Sebastian Blaes, Georg Martius
Abstract
Solving high-dimensional, continuous robotic tasks is a challenging optimization problem. Model-based methods that rely on zero-order optimizers like the crossentropy method (CEM) have so far shown strong performance and are considered state-of-the-art in the model-based reinforcement learning community. However, this success comes at the cost of high computational complexity, being therefore not suitable for real-time control. In this paper, we propose a technique to jointly optimize the trajectory and distill a policy, which is essential for fast execution in real robotic systems. Our method builds upon standard approaches, like guidance cost and dataset aggregation, and introduces a novel adaptive factor which prevents the optimizer from collapsing to the learner's behavior at the beginning of the training. The extracted policies reach unprecedented performance on challenging tasks like making a humanoid stand up and opening a door without reward shaping. Figure 1: Environments and exemplary behaviors of the learned policy using APEX. From left to right: FETCH PICK&PLACE (sparse reward), DOOR (sparse reward), and HUMANOID STANDUP. * equal contribution. We acknowledge the support from the German Federal Ministry of Education and Research (BMBF) through the Tübingen AI Center (FKZ: 01IS18039B) and from the Max Planck ETH Center for Learning Systems. Recently, approaches like simple point-to-point supervised training such as Behavioral Cloning (BC), or Generative Adversarial Network training (GAN) have been explored (Wang & Ba, 2020) for policy distillation from CEM, but only largely sub-optimal policies could be extracted. When the policy is used alone at test time and not in combination with the MPC-CEM optimizer, its performance drops
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4f3b064-c423-4135-83e3-9d0407fc618dCited by top-tier papers3
- On the Pitfalls of Heteroscedastic Uncertainty Estimation with Probabilistic Neural NetworksMaximilian Seitzer, Arash Tavakoli, Dimitrije Antic, Georg MartiusICLR 2022 · 122 citations
- Sparsely Changing Latent States for Prediction and Planning in Partially Observable DomainsChristian Gumbsch, Martin V. Butz, Georg MartiusNeurIPS 2021 · 30 citations
- Neuro-algorithmic Policies Enable Fast Combinatorial GeneralizationMarin Vlastelica P., Michal Rolínek, Georg MartiusICML 2021 · 17 citations
Builds on1
Related papers
- The Differentiable Cross-Entropy MethodBrandon Amos, Denis YaratsICML 2020 · 60 citations
- Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task LearnersMichal Nauman, Marek Cygan, Carmelo Sferrazza, Aviral Kumar et al.NeurIPS 2025 · 26 citations
- Evaluating Model-Based Planning and Planner Amortization for Continuous ControlArunkumar Byravan, Leonard Hasenclever, Piotr Trochim, Mehdi Mirza et al.ICLR 2022 · 18 citations
- Bridging Environments and Language with Rendering Functions and Vision-Language ModelsThéo Cachet, Christopher R. Dance, Olivier SigaudICML 2024 · 1 citation
- Scaffolding Dexterous Manipulation with Vision-Language ModelsVincent de Bakker, Joey Hejna, Tyler Ga Wei Lum, Onur Celik et al.NeurIPS 2025 · 14 citations
