SIFTER: Space-Efficient Value Iteration for Finite-Horizon MDPs
Konstantinos Skitsas, Ioannis G. Papageorgiou, Mohammad Sadegh Talebi, Vasiliki Kantere, Michael N. Katehakis, Panagiotis Karras
Abstract
Can we solve finite-horizon Markov decision processes (FHMDPs) while raising low memory requirements? Such models find application in many cases where a decision-making agent needs to act in a probabilistic environment, from resource management to medicine to service provisioning. However, computing optimal policies such an agent should follow by dynamic programming value iteration raises either prohibitive space complexity, or, in reverse, non-scalable time complexity requirements. This scalability question has been largely neglected. In this paper, we propose SIFTER (Space Efficient Finite Horizon MDPs), a suite of algorithms that achieve a golden middle between space and time requirements. Our former algorithm raises space complexity growing with the square root of the horizon's length without a time-complexity overhead, while the latter's space requirements depend only logarithmically in horizon length with a corresponding logarithmic time complexity overhead. A thorough experimental study under diverse settings confirms that SIFTER algorithms achieve the predicted gains, while approximation techniques do not achieve the same combination of time efficiency, space efficiency, and result quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- SIEVE: A Space-Efficient Algorithm for Viterbi DecodingMartino Ciaperoni, Aristides Gionis, Athanasios Katsamanis, Panagiotis KarrasSIGMOD 2022 · 3 citations
- Sketched Newton Value Iteration for Large-Scale Markov Decision ProcessesJinsong Liu, Chenghan Xie, Qi Deng, Dongdong Ge et al.AAAI 2024 · 1 citation
- Quantum Algorithms for Finite-horizon Markov Decision ProcessesBin Luo, Yuwen Huang, Jonathan Allcock, Xiaojun Lin et al.ICML 2025
- A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPsKihyuk Hong, Ambuj TewariICML 2025
- Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative ModelGen Li, Yuting Wei, Yuejie Chi, Yuantao Gu et al.NeurIPS 2020 · 159 citations
