The Differentiable Cross-Entropy Method
Brandon Amos, Denis Yarats
Abstract
We study the cross-entropy method (CEM) for the non-convex optimization of a continuous and parameterized objective function and introduce a differentiable variant that enables us to differentiate the output of CEM with respect to the objective function's parameters. In the machine learning setting this brings CEM inside of the end-to-end learning pipeline where this has otherwise been impossible. We show applications in a synthetic energy-based structured prediction task and in non-convex continuous control. In the control setting we show how to embed optimal action sequences into a lower-dimensional space. DCEM enables us to fine-tune CEM-based controllers with policy optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec4685db-26dd-416f-bb71-b41287b3eb93Cited by top-tier papers16
- CombOptNet: Fit the Right NP-Hard Problem by Learning Integer Programming ConstraintsAnselm Paulus, Michal Rolínek, Vít Musil, Brandon Amos et al.ICML 2021 · 73 citations
- Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud RegistrationHaobo Jiang, Yaqi Shen, Jin Xie, Jun Li et al.ICCV 2021 · 52 citations
- Control-Oriented Model-Based Reinforcement Learning with Implicit DifferentiationEvgenii Nikishin, Romina Abachi, Rishabh Agarwal, Pierre-Luc BaconAAAI 2022 · 47 citations
- Latent Skill Planning for Exploration and TransferKevin Xie, Homanga Bharadhwaj, Danijar Hafner, Animesh Garg et al.ICLR 2021 · 27 citations
- Iterative Amortized Policy OptimizationJoseph Marino, Alexandre Piché, Alessandro Davide Ialongo, Yisong YueNeurIPS 2021 · 27 citations
Builds on2
Related papers
- NOVAS: Non-convex Optimization via Adaptive Stochastic Search for End-to-end Learning and ControlIoannis Exarchos, Marcus Aloysius Pereira, Ziyi Wang, Evangelos A. TheodorouICLR 2021 · 4 citations
- Greedy Actor-Critic: A New Conditional Cross-Entropy Method for Policy ImprovementSamuel Neumann, Sungsu Lim, Ajin George Joseph, Yangchen Pan et al.ICLR 2023 · 3 citations
- Extracting Strong Policies for Robotics Tasks from Zero-Order Trajectory OptimizersCristina Pinneri, Shambhuraj Sawant, Sebastian Blaes, Georg MartiusICLR 2021 · 14 citations
- Efficient Training of Energy-Based Models Using Jarzynski EqualityDavide Carbone, Mengjian Hua, Simon Coste, Eric Vanden-EijndenNeurIPS 2023 · 21 citations
- Maximum Entropy Reinforcement Learning via Energy-Based Normalizing FlowChen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee et al.NeurIPS 2024 · 29 citations
