Search and Explore: Symbiotic Policy Synthesis in POMDPs
Roman Andriushchenko, Alexander Bork, Milan Ceska, Sebastian Junges, Joost-Pieter Katoen, Filip Macák
Abstract
Abstract This paper marries two state-of-the-art controller synthesis methods for partially observable Markov decision processes (POMDPs), a prominent model in sequential decision making under uncertainty. A central issue is to find a POMDP controller—that solely decides based on the observations seen so far—to achieve a total expected reward objective. As finding optimal controllers is undecidable, we concentrate on synthesising good finite-state controllers (FSCs). We do so by tightly integrating two modern, orthogonal methods for POMDP controller synthesis: a belief-based and an inductive approach. The former method obtains an FSC from a finite fragment of the so-called belief MDP, an MDP that keeps track of the probabilities of equally observable POMDP states. The latter is an inductive search technique over a set of FSCs, e.g., controllers with a fixed memory size. The key result of this paper is a symbiotic anytime algorithm that tightly integrates both approaches such that each profits from the controllers constructed by the other. Experimental results indicate a substantial improvement in the value of the controllers while significantly reducing the synthesis time and memory footprint.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c94e9174-0fc2-4bf3-9cf0-4bf65c30fce3Cited by top-tier papers2
- Constrained and Robust Policy Synthesis with Satisfiability-Modulo-Probabilistic-Model-CheckingLinus Heck, Filip Macák, Milan Ceska, Sebastian JungesAAAI 2026
- Synthesizing POMDP Policies: Sampling Meets Model-Checking via LearningDebraj Chakraborty, Anirban Majumdar, Prince Mathew, Sayan Mukherjee et al.CAV 2026
Builds on1
Related papers
- What Should Be Observed for Optimal Reward in POMDPs?Alyzia-Maria Konsta, Alberto Lluch-Lafuente, Christoph MathejaCAV 2024 · 2 citations
- Point-Based Methods for Model Checking in Partially Observable Markov Decision ProcessesMaxime Bouton, Jana Tumova, Mykel J. KochenderferAAAI 2020 · 32 citations
- Reference-Based POMDPsEdward Kim, Yohan Karunanayake, Hanna KurniawatiNeurIPS 2023 · 5 citations
- Revelations: A Decidable Class of POMDPs with Omega-Regular ObjectivesMarius Belly, Nathanaël Fijalkow, Hugo Gimbert, Florian Horn et al.AAAI 2025 · 5 citations
- Enforcing Almost-Sure Reachability in POMDPsSebastian Junges, Nils Jansen, Sanjit A. SeshiaCAV 2021 · 8 citations
