Search and Explore: Symbiotic Policy Synthesis in POMDPs
Roman Andriushchenko, Alexander Bork, Milan Ceska, Sebastian Junges, Joost-Pieter Katoen, Filip Macák
摘要
Abstract This paper marries two state-of-the-art controller synthesis methods for partially observable Markov decision processes (POMDPs), a prominent model in sequential decision making under uncertainty. A central issue is to find a POMDP controller—that solely decides based on the observations seen so far—to achieve a total expected reward objective. As finding optimal controllers is undecidable, we concentrate on synthesising good finite-state controllers (FSCs). We do so by tightly integrating two modern, orthogonal methods for POMDP controller synthesis: a belief-based and an inductive approach. The former method obtains an FSC from a finite fragment of the so-called belief MDP, an MDP that keeps track of the probabilities of equally observable POMDP states. The latter is an inductive search technique over a set of FSCs, e.g., controllers with a fixed memory size. The key result of this paper is a symbiotic anytime algorithm that tightly integrates both approaches such that each profits from the controllers constructed by the other. Experimental results indicate a substantial improvement in the value of the controllers while significantly reducing the synthesis time and memory footprint.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Constrained and Robust Policy Synthesis with Satisfiability-Modulo-Probabilistic-Model-CheckingLinus Heck, Filip Macák, Milan Ceska, Sebastian JungesAAAI 2026
- Synthesizing POMDP Policies: Sampling Meets Model-Checking via LearningDebraj Chakraborty, Anirban Majumdar, Prince Mathew, Sayan Mukherjee 等CAV 2026
它引用的顶会 Paper1
相关 Paper
- What Should Be Observed for Optimal Reward in POMDPs?Alyzia-Maria Konsta, Alberto Lluch-Lafuente, Christoph MathejaCAV 2024 · 被引用 2 次
- Point-Based Methods for Model Checking in Partially Observable Markov Decision ProcessesMaxime Bouton, Jana Tumova, Mykel J. KochenderferAAAI 2020 · 被引用 32 次
- Reference-Based POMDPsEdward Kim, Yohan Karunanayake, Hanna KurniawatiNeurIPS 2023 · 被引用 5 次
- Revelations: A Decidable Class of POMDPs with Omega-Regular ObjectivesMarius Belly, Nathanaël Fijalkow, Hugo Gimbert, Florian Horn 等AAAI 2025 · 被引用 5 次
- Enforcing Almost-Sure Reachability in POMDPsSebastian Junges, Nils Jansen, Sanjit A. SeshiaCAV 2021 · 被引用 8 次
