SPO: Sequential Monte Carlo Policy Optimisation
Matthew Macfarlane, Edan Toledo, Donal Byrne, Paul Duckworth, Alexandre Laterre
摘要
Leveraging planning during learning and decision-making is central to the long-term development of intelligent agents. Recent works have successfully combined tree-based search methods and self-play learning mechanisms to this end. However, these methods typically face scaling challenges due to the sequential nature of their search. While practical engineering solutions can partly overcome this, they often result in a negative impact on performance. In this paper, we introduce SPO: Sequential Monte Carlo Policy Optimisation, a model-based reinforcement learning algorithm grounded within the Expectation Maximisation (EM) framework. We show that SPO provides robust policy improvement and efficient scaling properties. The sample-based search makes it directly applicable to both discrete and continuous action spaces without modifications. We demonstrate statistically significant improvements in performance relative to model-free and model-based baselines across both continuous and discrete environments. Furthermore, the parallel nature of SPO's search enables effective utilisation of hardware accelerators, yielding favourable scaling laws.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Twice Sequential Monte Carlo for Tree SearchYaniv Oren, Joery de Vries, Pascal Van der Vaart, Matthijs T. J. Spaan 等ICML 2026 · 被引用 2 次
- Breaking the Performance Ceiling in Reinforcement Learning requires Inference StrategiesFélix Chalumeau, Daniel Rajaonarivonivelomanantsoa, Ruan John de Kock, Juan Claude Formanek 等NeurIPS 2025 · 被引用 2 次
- Trust-Region Twisted Policy ImprovementJoery A. de Vries, Jinke He, Yaniv Oren, Matthijs T. J. SpaanICML 2025
它引用的顶会 Paper17
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous ControlH. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark 等ICLR 2020 · 被引用 138 次
- Deep active inference agents using Monte-Carlo methodsZafeirios Fountas, Noor Sajid, Pedro A. M. Mediano, Karl J. FristonNeurIPS 2020 · 被引用 130 次
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 被引用 120 次
- Constrained Variational Policy Optimization for Safe Reinforcement LearningZuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Wei Liu 等ICML 2022 · 被引用 112 次
相关 Paper
- Making Better Decision by Directly Planning in Continuous ControlJinhua Zhu, Yue Wang, Lijun Wu, Tao Qin 等ICLR 2023
- Efficient Multi-agent Reinforcement Learning by PlanningQihan Liu, Jianing Ye, Xiaoteng Ma, Jun Yang 等ICLR 2024 · 被引用 18 次
- Parallelizing Model-based Reinforcement Learning Over the Sequence LengthZirui Wang, Yue Deng, Junfeng Long, Yin ZhangNeurIPS 2024 · 被引用 9 次
- Bayesian Optimized Monte Carlo PlanningJohn Mern, Anil Yildiz, Zachary Sunberg, Tapan Mukerji 等AAAI 2021 · 被引用 33 次
- Policy Gradient with Tree ExpansionGal Dalal, Assaf Hallak, Gugan Thoppe, Shie Mannor 等ICML 2025
