Optimizing Local Satisfaction of Long-Run Average Objectives in Markov Decision Processes
David Klaska, Antonín Kucera, Vojtech Kur, Vít Musil, Vojtech Rehák
2024Year
1Citations
Abstract
Long-run average optimization problems for Markov decision processes (MDPs) require constructing policies with optimal steady-state behavior, i.e., optimal limit frequency of visits to the states. However, such policies may suffer from local instability in the sense that the frequency of states visited in a bounded time horizon along a run differs significantly from the limit frequency. In this work, we propose an efficient algorithmic solution to this problem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Multiple Mean-Payoff Optimization Under Local Stability ConstraintsDavid Klaska, Antonín Kucera, Vojtech Kur, Vít Musil et al.AAAI 2025
- The Geometry of Memoryless Stochastic Policy Optimization in Infinite-Horizon POMDPsJohannes Müller, Guido MontúfarICLR 2022 · 9 citations
- Fast Mixing Steady-State Control in Markov Decision ProcessesFederico Corso, Marco Mussi, Alberto Maria MetelliICML 2026
- Efficient Planning in Large MDPs with Weak Linear Function ApproximationRoshan Shariff, Csaba SzepesváriNeurIPS 2020 · 23 citations
- Offline Actor-Critic for Average Reward MDPsWilliam G. Powell, Jeongyeol Kwon, Qiaomin Xie, Hanbaek LyuNeurIPS 2025
