The Geometry of Memoryless Stochastic Policy Optimization in Infinite-Horizon POMDPs
Johannes Müller, Guido Montúfar
摘要
We consider the problem of finding the best memoryless stochastic policy for an infinite-horizon partially observable Markov decision process (POMDP) with finite state and action spaces with respect to either the discounted or mean reward criterion. We show that the (discounted) state-action frequencies and the expected cumulative reward are rational functions of the policy, whereby the degree is determined by the degree of partial observability. We then describe the optimization problem as a linear optimization problem in the space of feasible state-action frequencies subject to polynomial constraints that we characterize explicitly. This allows us to address the combinatorial and geometric complexity of the optimization problem using recent tools from polynomial optimization. In particular, we estimate the number of critical points and use the polynomial programming description of reward maximization to solve a navigation problem in a grid world.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- On Shallow Planning Under Partial ObservabilityRandy Lefebvre, Audrey DurandAAAI 2025 · 被引用 2 次
- Geometric Policy Iteration for Markov Decision ProcessesYue Wu, Jesús A. De LoeraKDD 2022 · 被引用 1 次
- On Minimizing Adversarial Counterfactual Error in Adversarial Reinforcement LearningRoman Belaire, Arunesh Sinha, Pradeep VarakanthamICLR 2025
它引用的顶会 Paper3
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Reward is enough for convex MDPsTom Zahavy, Brendan O'Donoghue, Guillaume Desjardins, Satinder SinghNeurIPS 2021 · 被引用 96 次
- Pure and Spurious Critical Points: a Geometric Study of Linear NetworksMatthew Trager, Kathlén Kohn, Joan BrunaICLR 2020 · 被引用 41 次
相关 Paper
- The Value Function Semi-Algebraic Set in Partially Observable Markov Decision ProcessesRyan Anderson, Guido MontufarICML 2026
- Point-Based Methods for Model Checking in Partially Observable Markov Decision ProcessesMaxime Bouton, Jana Tumova, Mykel J. KochenderferAAAI 2020 · 被引用 32 次
- Reference-Based POMDPsEdward Kim, Yohan Karunanayake, Hanna KurniawatiNeurIPS 2023 · 被引用 5 次
- Reinforcement Learning from Partial Observation: Linear Function Approximation with Provable Sample EfficiencyQi Cai, Zhuoran Yang, Zhaoran WangICML 2022 · 被引用 17 次
- Offline Actor-Critic for Average Reward MDPsWilliam G. Powell, Jeongyeol Kwon, Qiaomin Xie, Hanbaek LyuNeurIPS 2025
