Lune

CAV2024Top-tier venue

What Should Be Observed for Optimal Reward in POMDPs?

Alyzia-Maria Konsta, Alberto Lluch-Lafuente, Christoph Matheja

2024Year
2Citations
1Top-tier citations

Abstract

Abstract Partially observable Markov Decision Processes (POMDPs) are a standard model for agents making decisions in uncertain environments. Most work on POMDPs focuses on synthesizing strategies based on the available capabilities. However, system designers can often control an agent’s observation capabilities, e.g. by placing or selecting sensors. This raises the question of how one should select an agent’s sensors cost-effectively such that it achieves the desired goals. In this paper, we study the noveloptimal observability problem(oop): Given a POMDP M\mathscr {M} M , how should one change M\mathscr {M} M ’s observation capabilities within a fixed budget such that its (minimal) expected reward remains below a given threshold? We show that the problem is undecidable in general and decidable when considering positional strategies only. We present two algorithms for a decidable fragment of theoop: one based on optimal strategies of M\mathscr {M} M ’s underlying Markov decision process and one based on parameter synthesis with SMT. We report promising results for variants of typical examples from the POMDP literature.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e89213b8-d191-46d0-a16a-a0bd647a039b

Cited by top-tier papers1

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines