Optimizing Local Satisfaction of Long-Run Average Objectives in Markov Decision Processes
David Klaska, Antonín Kucera, Vojtech Kur, Vít Musil, Vojtech Rehák
2024年份
1被引次数
摘要
Long-run average optimization problems for Markov decision processes (MDPs) require constructing policies with optimal steady-state behavior, i.e., optimal limit frequency of visits to the states. However, such policies may suffer from local instability in the sense that the frequency of states visited in a bounded time horizon along a run differs significantly from the limit frequency. In this work, we propose an efficient algorithmic solution to this problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Multiple Mean-Payoff Optimization Under Local Stability ConstraintsDavid Klaska, Antonín Kucera, Vojtech Kur, Vít Musil 等AAAI 2025
- The Geometry of Memoryless Stochastic Policy Optimization in Infinite-Horizon POMDPsJohannes Müller, Guido MontúfarICLR 2022 · 被引用 9 次
- Fast Mixing Steady-State Control in Markov Decision ProcessesFederico Corso, Marco Mussi, Alberto Maria MetelliICML 2026
- Efficient Planning in Large MDPs with Weak Linear Function ApproximationRoshan Shariff, Csaba SzepesváriNeurIPS 2020 · 被引用 23 次
- Offline Actor-Critic for Average Reward MDPsWilliam G. Powell, Jeongyeol Kwon, Qiaomin Xie, Hanbaek LyuNeurIPS 2025
