Critic Sequential Monte Carlo
Vasileios Lioutas, Jonathan Wilder Lavington, Justice Sefas, Matthew Niedoba, Yunpeng Liu, Berend Zwartsenberg, Setareh Dabiri, Frank Wood, Adam Scibior
Abstract
We introduce CriticSMC, a new algorithm for planning as inference built from a composition of sequential Monte Carlo with learned Soft-Q function heuristic factors. These heuristic factors, obtained from parametric approximations of the marginal likelihood ahead, more effectively guide SMC towards the desired target distribution, which is particularly helpful for planning in environments with hard constraints placed sparsely in time. Compared with previous work, we modify the placement of such heuristic factors, which allows us to cheaply propose and evaluate large numbers of putative action particles, greatly increasing inference and planning efficiency. CriticSMC is compatible with informative priors, whose density function need not be known, and can be used as a model-free control algorithm. Our experiments on collision avoidance in a high-dimensional simulated driving task show that CriticSMC significantly reduces collision rates at a low computational cost while maintaining realism and diversity of driving behaviors across vehicles and environment scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e50dc5f-2915-46b7-93b7-a6233a3db89cCited by top-tier papers7
- Probabilistic Inference in Language Models via Twisted Sequential Monte CarloStephen Zhao, Rob Brekelmans, Alireza Makhzani, Roger Baker GrosseICML 2024 · 61 citations
- SPO: Sequential Monte Carlo Policy OptimisationMatthew Macfarlane, Edan Toledo, Donal Byrne, Paul Duckworth et al.NeurIPS 2024 · 8 citations
- NAS-X: Neural Adaptive Smoothing via TwistingDieterich Lawson, Michael Li, Scott W. LindermanNeurIPS 2023 · 3 citations
- Mitigating Premature Exploitation in Particle-based Monte Carlo for Inference-Time ScalingGiorgio Giannone, Guangxuan Xu, Nikhil Nayak, Rohan Awhad et al.ICML 2026 · 2 citations
- Twice Sequential Monte Carlo for Tree SearchYaniv Oren, Joery de Vries, Pascal Van der Vaart, Matthijs T. J. Spaan et al.ICML 2026 · 2 citations
Builds on4
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement LearningAviral Kumar, Rishabh Agarwal, Dibya Ghosh, Sergey LevineICLR 2021 · 155 citations
- CAQL: Continuous Action Q-LearningMoonkyung Ryu, Yinlam Chow, Ross Anderson, Christian Tjandraatmadja et al.ICLR 2020 · 50 citations
- Variational Inference for Sequential Data with Future Likelihood EstimatesGeon-Hyeong Kim, Youngsoo Jang, Hongseok Yang, Kee-Eung KimICML 2020 · 6 citations
Related papers
- Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic LearningHaque Ishfaq, Guangyuan Wang, Sami Nur Islam, Doina PrecupICLR 2025
- Off-Policy Safe Reinforcement Learning with Cost-Constrained Optimistic ExplorationGuopeng Li, Matthijs T. J. Spaan, Julian F. P. KooijICLR 2026
- Monte Carlo Augmented Actor-Critic for Sparse Reward Deep Reinforcement Learning from Suboptimal DemonstrationsAlbert Wilcox, Ashwin Balakrishna, Jules Dedieu, Wyame Benslimane et al.NeurIPS 2022 · 28 citations
- A Probabilistic Framework for LLM-Based Model DiscoveryStefan Wahl, Raphaela Schenk, Ali Farnoud, Jakob Macke et al.ICML 2026 · 7 citations
- Trust-Region Twisted Policy ImprovementJoery A. de Vries, Jinke He, Yaniv Oren, Matthijs T. J. SpaanICML 2025
