The Differentiable Cross-Entropy Method
Brandon Amos, Denis Yarats
摘要
We study the cross-entropy method (CEM) for the non-convex optimization of a continuous and parameterized objective function and introduce a differentiable variant that enables us to differentiate the output of CEM with respect to the objective function's parameters. In the machine learning setting this brings CEM inside of the end-to-end learning pipeline where this has otherwise been impossible. We show applications in a synthetic energy-based structured prediction task and in non-convex continuous control. In the control setting we show how to embed optimal action sequences into a lower-dimensional space. DCEM enables us to fine-tune CEM-based controllers with policy optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- CombOptNet: Fit the Right NP-Hard Problem by Learning Integer Programming ConstraintsAnselm Paulus, Michal Rolínek, Vít Musil, Brandon Amos 等ICML 2021 · 被引用 73 次
- Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud RegistrationHaobo Jiang, Yaqi Shen, Jin Xie, Jun Li 等ICCV 2021 · 被引用 52 次
- Control-Oriented Model-Based Reinforcement Learning with Implicit DifferentiationEvgenii Nikishin, Romina Abachi, Rishabh Agarwal, Pierre-Luc BaconAAAI 2022 · 被引用 47 次
- Latent Skill Planning for Exploration and TransferKevin Xie, Homanga Bharadhwaj, Danijar Hafner, Animesh Garg 等ICLR 2021 · 被引用 27 次
- Iterative Amortized Policy OptimizationJoseph Marino, Alexandre Piché, Alessandro Davide Ialongo, Yisong YueNeurIPS 2021 · 被引用 27 次
它引用的顶会 Paper2
相关 Paper
- NOVAS: Non-convex Optimization via Adaptive Stochastic Search for End-to-end Learning and ControlIoannis Exarchos, Marcus Aloysius Pereira, Ziyi Wang, Evangelos A. TheodorouICLR 2021 · 被引用 4 次
- Greedy Actor-Critic: A New Conditional Cross-Entropy Method for Policy ImprovementSamuel Neumann, Sungsu Lim, Ajin George Joseph, Yangchen Pan 等ICLR 2023 · 被引用 3 次
- Extracting Strong Policies for Robotics Tasks from Zero-Order Trajectory OptimizersCristina Pinneri, Shambhuraj Sawant, Sebastian Blaes, Georg MartiusICLR 2021 · 被引用 14 次
- Efficient Training of Energy-Based Models Using Jarzynski EqualityDavide Carbone, Mengjian Hua, Simon Coste, Eric Vanden-EijndenNeurIPS 2023 · 被引用 21 次
- Maximum Entropy Reinforcement Learning via Energy-Based Normalizing FlowChen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee 等NeurIPS 2024 · 被引用 29 次
