A Simple Decentralized Cross-Entropy Method
Zichen Zhang, Jun Jin, Martin Jägersand, Jun Luo, Dale Schuurmans
Abstract
Cross-Entropy Method (CEM) is commonly used for planning in model-based reinforcement learning (MBRL) where a centralized approach is typically utilized to update the sampling distribution based on only the top- operation's results on samples. In this paper, we show that such a centralized approach makes CEM vulnerable to local optima, thus impairing its sample efficiency. To tackle this issue, we propose Decentralized CEM (DecentCEM), a simple but effective improvement over classical CEM, by using an ensemble of CEM instances running independently from one another, and each performing a local improvement of its own sampling distribution. We provide both theoretical and empirical analysis to demonstrate the effectiveness of this simple decentralized approach. We empirically show that, compared to the classical centralized approach using either a single or even a mixture of Gaussian distributions, our DecentCEM finds the global optimum much more consistently thus improves the sample efficiency. Furthermore, we plug in our DecentCEM in the planning problem of MBRL, and evaluate our approach in several continuous control environments, with comparison to the state-of-art CEM based MBRL approaches (PETS and POPLIN). Results show sample efficiency improvement by simply replacing the classical CEM module with our DecentCEM module, while only sacrificing a reasonable amount of computational cost. Lastly, we conduct ablation studies for more in-depth analysis. Code is available at https://github.com/vincentzhang/decentCEM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47662cf7-08d5-4fd0-bbe3-cb8d713273a4Cited by top-tier papers2
- Self-Labeling the Job Shop Scheduling ProblemAndrea Corsini, Angelo Porrello, Simone Calderara, Mauro Dell'AmicoNeurIPS 2024 · 39 citations
- BaB-ND: Long-Horizon Motion Planning with Branch-and-Bound and Neural DynamicsKeyi Shen, Jiangwei Yu, Jose A. Barreiros, Huan Zhang et al.ICLR 2025
Builds on3
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 164 citations
- The Differentiable Cross-Entropy MethodBrandon Amos, Denis YaratsICML 2020 · 60 citations
- An Efficient Asynchronous Method for Integrating Evolutionary and Gradient-based Policy SearchKyunghyun Lee, Byeong-Uk Lee, Ukcheol Shin, In So KweonNeurIPS 2020 · 24 citations
Related papers
- Making Better Decision by Directly Planning in Continuous ControlJinhua Zhu, Yue Wang, Lijun Wu, Tao Qin et al.ICLR 2023
- Double Buffers CEM-TD3: More Efficient Evolution and Richer ExplorationSheng Zhu, Chun Shen, Shuai Lü, Junhong Wu et al.AAAI 2024 · 3 citations
- Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud RegistrationHaobo Jiang, Yaqi Shen, Jin Xie, Jun Li et al.ICCV 2021 · 52 citations
- Sample-Efficient Reinforcement Learning via Conservative Model-Based Actor-CriticZhihai Wang, Jie Wang, Qi Zhou, Bin Li et al.AAAI 2022 · 38 citations
- SUMO: Search-Based Uncertainty Estimation for Model-Based Offline Reinforcement LearningZhongjian Qiao, Jiafei Lyu, Kechen Jiao, Qi Liu et al.AAAI 2025 · 12 citations
