A Simple Decentralized Cross-Entropy Method
Zichen Zhang, Jun Jin, Martin Jägersand, Jun Luo, Dale Schuurmans
摘要
Cross-Entropy Method (CEM) is commonly used for planning in model-based reinforcement learning (MBRL) where a centralized approach is typically utilized to update the sampling distribution based on only the top- operation's results on samples. In this paper, we show that such a centralized approach makes CEM vulnerable to local optima, thus impairing its sample efficiency. To tackle this issue, we propose Decentralized CEM (DecentCEM), a simple but effective improvement over classical CEM, by using an ensemble of CEM instances running independently from one another, and each performing a local improvement of its own sampling distribution. We provide both theoretical and empirical analysis to demonstrate the effectiveness of this simple decentralized approach. We empirically show that, compared to the classical centralized approach using either a single or even a mixture of Gaussian distributions, our DecentCEM finds the global optimum much more consistently thus improves the sample efficiency. Furthermore, we plug in our DecentCEM in the planning problem of MBRL, and evaluate our approach in several continuous control environments, with comparison to the state-of-art CEM based MBRL approaches (PETS and POPLIN). Results show sample efficiency improvement by simply replacing the classical CEM module with our DecentCEM module, while only sacrificing a reasonable amount of computational cost. Lastly, we conduct ablation studies for more in-depth analysis. Code is available at https://github.com/vincentzhang/decentCEM
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Self-Labeling the Job Shop Scheduling ProblemAndrea Corsini, Angelo Porrello, Simone Calderara, Mauro Dell'AmicoNeurIPS 2024 · 被引用 39 次
- BaB-ND: Long-Horizon Motion Planning with Branch-and-Bound and Neural DynamicsKeyi Shen, Jiangwei Yu, Jose A. Barreiros, Huan Zhang 等ICLR 2025
它引用的顶会 Paper3
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 被引用 164 次
- The Differentiable Cross-Entropy MethodBrandon Amos, Denis YaratsICML 2020 · 被引用 60 次
- An Efficient Asynchronous Method for Integrating Evolutionary and Gradient-based Policy SearchKyunghyun Lee, Byeong-Uk Lee, Ukcheol Shin, In So KweonNeurIPS 2020 · 被引用 24 次
相关 Paper
- Making Better Decision by Directly Planning in Continuous ControlJinhua Zhu, Yue Wang, Lijun Wu, Tao Qin 等ICLR 2023
- Double Buffers CEM-TD3: More Efficient Evolution and Richer ExplorationSheng Zhu, Chun Shen, Shuai Lü, Junhong Wu 等AAAI 2024 · 被引用 3 次
- Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud RegistrationHaobo Jiang, Yaqi Shen, Jin Xie, Jun Li 等ICCV 2021 · 被引用 52 次
- Sample-Efficient Reinforcement Learning via Conservative Model-Based Actor-CriticZhihai Wang, Jie Wang, Qi Zhou, Bin Li 等AAAI 2022 · 被引用 38 次
- SUMO: Search-Based Uncertainty Estimation for Model-Based Offline Reinforcement LearningZhongjian Qiao, Jiafei Lyu, Kechen Jiao, Qi Liu 等AAAI 2025 · 被引用 12 次
